WO2025007360A1 - 编解码方法、码流、编码器、解码器以及存储介质 - Google Patents

编解码方法、码流、编码器、解码器以及存储介质 Download PDF

Info

Publication number
WO2025007360A1
WO2025007360A1 PCT/CN2023/106200 CN2023106200W WO2025007360A1 WO 2025007360 A1 WO2025007360 A1 WO 2025007360A1 CN 2023106200 W CN2023106200 W CN 2023106200W WO 2025007360 A1 WO2025007360 A1 WO 2025007360A1
Authority
WO
WIPO (PCT)
Prior art keywords
current layer
value
mode
node
nodes
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/106200
Other languages
English (en)
French (fr)
Inventor
孙泽星
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Guangdong Oppo Mobile Telecommunications Corp Ltd
Original Assignee
Guangdong Oppo Mobile Telecommunications Corp Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Guangdong Oppo Mobile Telecommunications Corp Ltd filed Critical Guangdong Oppo Mobile Telecommunications Corp Ltd
Priority to CN202380099308.3A priority Critical patent/CN121359460A/zh
Priority to PCT/CN2023/106200 priority patent/WO2025007360A1/zh
Publication of WO2025007360A1 publication Critical patent/WO2025007360A1/zh
Priority to US19/423,905 priority patent/US20260113436A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/103Selection of coding mode or of prediction mode
    • H04N19/107Selection of coding mode or of prediction mode between spatial and temporal predictive coding, e.g. picture refresh
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/187Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being a scalable video layer
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/189Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the adaptation method, adaptation tool or adaptation type used for the adaptive coding
    • H04N19/196Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the adaptation method, adaptation tool or adaptation type used for the adaptive coding being specially adapted for the computation of encoding parameters, e.g. by averaging previously computed encoding parameters
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/30Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using hierarchical techniques, e.g. scalability
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/597Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding specially adapted for multi-view video sequence encoding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/70Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/90Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using coding techniques not provided for in groups H04N19/10-H04N19/85, e.g. fractals
    • H04N19/96Tree coding, e.g. quad-tree coding

Definitions

  • the embodiments of the present application relate to the field of point cloud encoding and decoding technology, and in particular, to an encoding and decoding method, a bit stream, an encoder, a decoder, and a storage medium.
  • G-PCC Geometry-based Point Cloud Compression
  • V-PCC Video-based Point Cloud Compression
  • MPEG Moving Picture Experts Group
  • attribute information encoding is mainly aimed at the encoding of color information.
  • color information encoding there are mainly two transformation methods.
  • One is the distance-based lifting transformation that relies on the level of detail (LOD) division, and the other is the direct region adaptive hierarchical transformation (RAHT).
  • LOD level of detail
  • RAHT direct region adaptive hierarchical transformation
  • the attribute encoding method of the entire sequence can be determined by the attribute parameter set (APS), for example, whether the entire sequence is encoded using RAHT transformation, prediction transformation, or lifting transformation.
  • APS attribute parameter set
  • this attribute encoding scheme does not fully consider the distribution of alternating components (Alternating Crrent) of different RAHT layers, resulting in low encoding efficiency of point cloud attributes.
  • the embodiments of the present application provide a coding and decoding method, a bit stream, an encoder, a decoder and a storage medium, which can improve the coding efficiency of point cloud attributes, thereby improving the coding and decoding performance of the point cloud.
  • an embodiment of the present application provides a decoding method, which is applied to a decoder, and the method includes:
  • the bitstream is parsed to determine the first syntax identification information
  • the first syntax identification information indicates that the current layer allows adaptive selection of an inter-frame prediction mode and/or an intra-frame prediction mode, parsing a bitstream to determine a target decoding mode of the current layer;
  • Attribute decoding is performed on the nodes in the current layer according to the target decoding mode to determine attribute reconstruction values of the nodes in the current layer.
  • an embodiment of the present application provides an encoding method, which is applied to an encoder, and the method includes:
  • determining a target coding mode of the current layer In a case where it is determined that a node of the current layer allows attribute prediction, and the current layer allows adaptive selection of an inter-frame prediction mode and/or an intra-frame prediction mode, determining a target coding mode of the current layer;
  • Attribute encoding is performed on the nodes in the current layer according to the target encoding mode to determine attribute reconstruction values of the nodes in the current layer.
  • an embodiment of the present application provides a code stream, which is generated by bit encoding according to information to be encoded; wherein the information to be encoded includes at least one of the following: a value of the first grammar identification information, a value of the second grammar identification information, a value of the third grammar identification information, a value of the fourth grammar identification information, a value of the fifth grammar identification information, a weight index value corresponding to the node in the current layer, and a second coefficient quantization residual value of the node in the current layer;
  • the first grammar identification information is used to indicate whether the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode
  • the second grammar identification information is used to indicate the target coding mode of the current layer
  • the value of the third grammar identification information is used to indicate the number of layers included in the current sequence where the current layer is located
  • the value of the fourth grammar identification information is used to indicate whether the nodes of the current layer are allowed to perform inter-frame prediction
  • the fifth grammar identification information is used to indicate whether the nodes of the current layer are allowed to perform intra-frame prediction
  • the sixth grammar identification information is used to indicate that the nodes in the current layer adopt a regional adaptive layered inter-frame transform mode
  • the weight index value is used to indicate the index value corresponding to the target weight combination corresponding to the nodes in the current layer in the preset weight table.
  • an embodiment of the present application provides a decoder, the decoder comprising a first determining unit and a decoding unit; wherein:
  • the first determining part is used to parse the bitstream and determine the first syntax identification information when it is determined that the node of the current layer allows attribute prediction; and parse the bitstream and determine the target decoding mode of the current layer when the first syntax identification information indicates that the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode;
  • the decoding part is used to perform attribute decoding on the nodes in the current layer according to the target decoding mode, and determine the attribute reconstruction values of the nodes in the current layer.
  • an embodiment of the present application provides a decoder, the decoder comprising a first memory and a first processor; wherein:
  • a first memory for storing a computer program that can be run on the first processor
  • the first processor is configured to execute the method according to the first aspect when running a computer program.
  • an embodiment of the present application provides an encoder, the encoder comprising a second determining unit and an encoding unit; wherein,
  • the second determination part is used to determine the target coding mode of the current layer and determine the first syntax identification information when it is determined that the node of the current layer allows attribute prediction and the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode; the first syntax identification information is used to indicate whether the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode;
  • the encoding part is used to perform attribute encoding on the nodes in the current layer according to the target encoding mode, and determine the attribute reconstruction values of the nodes in the current layer.
  • an encoder comprising a second memory and a second processor; wherein:
  • a second memory for storing a computer program that can be run on a second processor
  • the second processor is used to execute the method described in the second aspect when running the computer program.
  • an embodiment of the present application provides a computer-readable storage medium, which stores a computer program.
  • the computer program When executed, it implements the method described in the first aspect, or implements the method described in the second aspect.
  • the embodiment of the present application provides a coding and decoding method, a code stream, an encoder, a decoder and a storage medium.
  • the code stream is parsed to determine the first syntax identification information; when the first syntax identification information indicates that the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode, the code stream is parsed to determine the target decoding mode of the current layer; the nodes in the current layer are attribute-decoded according to the target decoding mode to determine the attribute reconstruction values of the nodes in the current layer.
  • the target coding mode of the current layer is determined, and the first syntax identification information is determined; the first syntax identification information is used to indicate whether the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode; the encoding part is used to attribute-encode the nodes in the current layer according to the target coding mode to determine the attribute reconstruction values of the nodes in the current layer.
  • the encoding end when performing attribute encoding for each layer, can adaptively select the target coding mode of each slice, and pass the target coding mode to the decoding end, so that the decoding end uses the parsed target decoding mode to reconstruct the attributes of the point cloud, thereby improving the encoding and decoding efficiency of the point cloud attributes, and then improving the encoding and decoding performance of the point cloud.
  • FIG1A is a schematic diagram of a three-dimensional point cloud image
  • FIG1B is a partial enlarged view of a three-dimensional point cloud image
  • FIG2A is a schematic diagram of six viewing angles of a point cloud image
  • FIG2B is a schematic diagram of a data storage format corresponding to a point cloud image
  • FIG3 is a schematic diagram of a network architecture for point cloud encoding and decoding
  • FIG4A is a schematic diagram of a composition framework of a G-PCC encoder
  • FIG4B is a schematic diagram of a composition framework of a G-PCC decoder
  • FIG5A is a schematic diagram of a low plane position in the Z-axis direction
  • FIG5B is a schematic diagram of a high plane position in the Z-axis direction
  • FIG6 is a schematic diagram of a node encoding sequence
  • FIG. 7A is a schematic diagram of a plane identification information
  • FIG7B is a schematic diagram of another type of planar identification information
  • FIG8 is a schematic diagram of sibling nodes of a current node
  • FIG9 is a schematic diagram of the intersection of a laser radar and a node
  • FIG10 is a schematic diagram of neighborhood nodes at the same partition depth and the same coordinates
  • FIG11 is a schematic diagram of a current node being located at a low plane position of a parent node
  • FIG12 is a schematic diagram of a high plane position of a current node located at a parent node
  • FIG13 is a schematic diagram of predictive coding of planar position information of a laser radar point cloud
  • FIG14 is a schematic diagram of IDCM encoding
  • FIG15 is a schematic diagram of coordinate transformation of a rotating laser radar to obtain a point cloud
  • FIG16 is a schematic diagram of predictive coding in the X-axis or Y-axis direction
  • FIG17A is a schematic diagram showing an angle of the Y plane predicted by a horizontal azimuth angle
  • FIG17B is a schematic diagram showing an angle of predicting the X-plane by using a horizontal azimuth angle
  • FIG18 is another schematic diagram of predictive coding in the X-axis or Y-axis direction
  • FIG19A is a schematic diagram of three intersection points included in a sub-block
  • FIG19B is a schematic diagram of a triangular facet set fitted using three intersection points
  • FIG19C is a schematic diagram of upsampling of a triangular face set
  • FIG20 is a schematic diagram of a distance-based LOD construction process
  • FIG21 is a schematic diagram of a visualization result of a LOD generation process
  • FIG22 is a schematic diagram of an encoding process for attribute prediction
  • FIG. 23 is a schematic diagram of the composition of a pyramid structure
  • FIG. 24 is a schematic diagram showing the composition of another pyramid structure
  • FIG25 is a schematic diagram of an LOD structure for inter-layer nearest neighbor search
  • FIG26 is a schematic diagram of a nearest neighbor search structure based on spatial relationship
  • FIG27A is a schematic diagram of a coplanar spatial relationship
  • FIG27B is a schematic diagram of a coplanar and colinear spatial relationship
  • FIG27C is a schematic diagram of a spatial relationship of coplanarity, colinearity and copointness
  • FIG28 is a schematic diagram of inter-layer prediction based on fast search
  • FIG29 is a schematic diagram of a LOD structure for nearest neighbor search within an attribute layer
  • FIG30 is a schematic diagram of intra-layer prediction based on fast search
  • FIG31 is a schematic diagram of a block-based neighborhood search structure
  • FIG32 is a schematic diagram of a coding process of a lifting transformation
  • FIG33 is a schematic diagram of a RAHT transformation structure
  • FIG34 is a schematic diagram of a RAHT transformation process along the x, y, and z directions;
  • FIG35A is a schematic diagram of a RAHT forward transformation process
  • FIG35B is a schematic diagram of a RAHT inverse transformation process
  • FIG36 is a schematic diagram of a flow chart of a decoding method provided in an embodiment of the present application.
  • FIG37 is a schematic diagram of the structure of an attribute coding block
  • FIG38 is a schematic diagram of the overall process of RAHT attribute prediction transform coding
  • FIG39 is a schematic diagram of a neighborhood prediction relationship of a current block
  • FIG40 is a schematic diagram of a calculation process of an attribute transformation coefficient
  • FIG41 is a schematic diagram of the structure of a RAHT attribute inter-frame prediction coding
  • FIG42 is a schematic diagram of a flow chart of an encoding method provided in an embodiment of the present application.
  • FIG43 is a schematic diagram of an attribute coding layer
  • FIG44 is a schematic diagram of the composition structure of a decoder provided in an embodiment of the present application.
  • FIG45 is a schematic diagram of a specific hardware structure of a decoder provided in an embodiment of the present application.
  • FIG46 is a schematic diagram of the composition structure of an encoder provided in an embodiment of the present application.
  • FIG47 is a schematic diagram of a specific hardware structure of an encoder provided in an embodiment of the present application.
  • Figure 48 is a schematic diagram of the composition structure of a coding and decoding system provided in an embodiment of the present application.
  • first ⁇ second ⁇ third involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that “first ⁇ second ⁇ third” can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
  • Point Cloud is a three-dimensional representation of the surface of an object.
  • Point Cloud (data) on the surface of an object can be collected through acquisition equipment such as photoelectric radar, lidar, laser scanner, and multi-view camera.
  • a point cloud is a set of irregularly distributed discrete points in space that express the spatial structure and surface properties of a three-dimensional object or scene.
  • FIG1A shows a three-dimensional point cloud image
  • FIG1B shows a partial magnified view of the three-dimensional point cloud image. It can be seen that the point cloud surface is composed of densely distributed points.
  • Two-dimensional images have information expressed at each pixel point, and the distribution is regular, so there is no need to record its position information additionally; however, the distribution of points in point clouds in three-dimensional space is random and irregular, so it is necessary to record the position of each point in space in order to fully express a point cloud.
  • each position in the acquisition process has corresponding attribute information, usually RGB color values, and the color value reflects the color of the object; for point clouds, in addition to color information, the attribute information corresponding to each point is also commonly the reflectance value, which reflects the surface material of the object. Therefore, point cloud data usually includes the position information of the point and the attribute information of the point. Among them, the position information of the point can also be called the geometric information of the point.
  • the geometric information of the point can be the three-dimensional coordinate information of the point (x, y, z).
  • the attribute information of the point can include color information and/or reflectivity, etc.
  • reflectivity can be one-dimensional reflectivity information (r); color information can be information on any color space, or color information can also be three-dimensional color information, such as RGB information.
  • R represents red (Red, R)
  • G represents green (Green, G)
  • B blue (Blue, B).
  • the color information may be luminance and chrominance (YCbCr, YUV) information, where Y represents brightness (Luma), Cb (U) represents blue color difference, and Cr (V) represents red color difference.
  • the points in the point cloud may include the three-dimensional coordinate information of the points and the reflectivity value of the points.
  • the points in the point cloud may include the three-dimensional coordinate information of the points and the three-dimensional color information of the points.
  • a point cloud obtained by combining the principles of laser measurement and photogrammetry may include the three-dimensional coordinate information of the points, the reflectivity value of the points and the three-dimensional color information of the points.
  • Figure 2A and 2B a point cloud image and its corresponding data storage format are shown.
  • Figure 2A provides six viewing angles of the point cloud image
  • Figure 2B consists of a file header information part and a data part.
  • the header information includes the data format, data representation type, the total number of point cloud points, and the content represented by the point cloud.
  • the point cloud is in the ".ply" format, represented by ASCII code, with a total number of 207242 points, and each point has three-dimensional coordinate information (x, y, z) and three-dimensional color information (r, g, b).
  • Point clouds can be divided into the following categories according to the way they are obtained:
  • Static point cloud the object is stationary, and the device that obtains the point cloud is also stationary;
  • Dynamic point cloud The object is moving, but the device that obtains the point cloud is stationary;
  • Dynamic point cloud acquisition The device used to acquire the point cloud is in motion.
  • point clouds can be divided into two categories according to their usage:
  • Category 1 Machine perception point cloud, which can be used in autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, disaster relief robots, etc.
  • Category 2 Point cloud perceived by the human eye, which can be used in point cloud application scenarios such as digital cultural heritage, free viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.
  • Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes. Point clouds are obtained by directly sampling real objects, so they can provide a strong sense of reality while ensuring accuracy. Therefore, they are widely used, including virtual reality games, computer-aided design, geographic information systems, automatic navigation systems, digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive remote presentation, and three-dimensional reconstruction of biological tissues and organs.
  • Point clouds can be collected mainly through the following methods: computer generation, 3D laser scanning, 3D photogrammetry, etc.
  • Computers can generate point clouds of virtual three-dimensional objects and scenes; 3D laser scanning can obtain point clouds of static real-world three-dimensional objects or scenes, and can obtain millions of point clouds per second; 3D photogrammetry can obtain point clouds of dynamic real-world three-dimensional objects or scenes, and can obtain tens of millions of point clouds per second.
  • 3D photogrammetry can obtain point clouds of dynamic real-world three-dimensional objects or scenes, and can obtain tens of millions of point clouds per second.
  • the number of points in each point cloud frame is 700,000, and each point has coordinate information xyz (float) and color information RGB (uchar).
  • point cloud compression has become a key issue in promoting the development of the point cloud industry.
  • the point cloud is a collection of massive points, storing the point cloud will not only consume a lot of memory, but also be inconvenient for transmission. There is also not enough bandwidth to support direct transmission of the point cloud at the network layer without compression. Therefore, the point cloud needs to be compressed.
  • the point cloud coding framework that can compress point clouds can be the geometry-based point cloud compression (G-PCC) codec framework provided by the Moving Picture Experts Group (MPEG) or
  • the video-based point cloud compression (V-PCC) codec framework can also be the AVS-PCC codec framework provided by AVS.
  • the G-PCC codec framework can be used to compress the first type of static point cloud and the third type of dynamically acquired point cloud, which can be based on the point cloud compression test platform (Test Model Compression 13, TMC13).
  • the V-PCC codec framework can be used to compress the second type of dynamic point cloud, which can be based on the point cloud compression test platform (Test Model Compression 2, TMC2). Therefore, the G-PCC codec framework is also called the point cloud codec TMC13, and the V-PCC codec framework is also called the point cloud codec TMC2.
  • FIG3 is a schematic diagram of a network architecture of a point cloud encoding and decoding provided by the embodiment of the present application.
  • the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication network 01.
  • the electronic device can be various types of devices with point cloud encoding and decoding functions.
  • the electronic device can include a mobile phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensor device, a server, etc., which is not limited by the embodiment of the present application.
  • the decoder or encoder in the embodiment of the present application can be the above-mentioned electronic device.
  • the electronic device in the embodiment of the present application has a point cloud encoding and decoding function, generally including a point cloud encoder (ie, encoder) and a point cloud decoder (ie, decoder).
  • a point cloud encoder ie, encoder
  • a point cloud decoder ie, decoder
  • the point cloud data is first divided into multiple slices by slice division.
  • the geometric information of the point cloud and the attribute information corresponding to each point are encoded separately.
  • FIG4A shows a schematic diagram of the composition framework of a G-PCC encoder.
  • the geometric information is transformed so that all point clouds are contained in a bounding box, and then quantized.
  • This step of quantization mainly plays a role in scaling. Due to the quantization rounding, the geometric information of a part of the point cloud is the same, so whether to remove duplicate points is determined based on parameters.
  • the process of quantization and removal of duplicate points is also called voxelization.
  • the bounding box is divided into octrees or a prediction tree is constructed.
  • arithmetic coding is performed on the points in the leaf nodes of the division to generate a binary geometric bit stream; or, arithmetic coding is performed on the intersection points (Vertex) generated by the division (surface fitting is performed based on the intersection points) to generate a binary geometric bit stream.
  • color conversion is required first to convert the color information (i.e., attribute information) from the RGB color space to the YUV color space. Then, the point cloud is recolored using the reconstructed geometric information so that the uncoded attribute information corresponds to the reconstructed geometric information. Attribute encoding is mainly performed on color information.
  • FIG4B shows a schematic diagram of the composition framework of a G-PCC decoder.
  • the geometric bit stream and the attribute bit stream in the binary bit stream are first decoded independently.
  • the geometric information of the point cloud is obtained through arithmetic decoding-reconstruction of the octree/reconstruction of the prediction tree-reconstruction of the geometry-coordinate inverse conversion;
  • the attribute information of the point cloud is obtained through arithmetic decoding-inverse quantization-LOD partitioning/RAHT-color inverse conversion, and the point cloud data to be encoded (i.e., the output point cloud) is restored based on the geometric information and attribute information.
  • the current geometric coding of G-PCC can be divided into octree-based geometric coding (marked by a dotted box) and prediction tree-based geometric coding (marked by a dotted box).
  • the octree-based geometry encoding includes: first, coordinate transformation of the geometric information so that all point clouds are contained in a bounding box. Then quantization is performed. This step of quantization mainly plays a role of scaling. Due to the quantization rounding, the geometric information of some points is the same. The parameters are used to decide whether to remove duplicate points. The process of quantization and removal of duplicate points is also called voxelization. Next, the bounding box is continuously divided into trees (such as octrees, quadtrees, binary trees, etc.) in the order of breadth-first traversal, and the placeholder code of each node is encoded.
  • trees such as octrees, quadtrees, binary trees, etc.
  • a company proposed an implicit geometry partitioning method.
  • the bounding box of the point cloud is calculated. Assume that dx > dy > dz , the bounding box corresponds to a cuboid.
  • binary tree partitioning will be performed based on the x-axis to obtain two child nodes.
  • quadtree partitioning will be performed based on the x- and y-axes to obtain four child nodes.
  • octree partitioning will be performed until the leaf node obtained by partitioning is a 1 ⁇ 1 ⁇ 1 unit cube.
  • K indicates the maximum number of binary tree/quadtree partitions before octree partitioning
  • M is used to indicate that the minimum block side length corresponding to binary tree/quadtree partitioning is 2M .
  • the reason why parameters K and M meet the above conditions is that in the current G-PCC geometric implicit partitioning process, the priority of the partitioning method is binary tree, quadtree and octree. When the node block size does not meet the binary tree/quadtree conditions, the node will be octreeed.
  • the octree is divided until the minimum unit of leaf nodes is 1 ⁇ 1 ⁇ 1.
  • the geometric information coding mode based on the octree can effectively encode the geometric information of the point cloud by utilizing the correlation between adjacent points in space.
  • the coding efficiency of the point cloud geometric information can be further improved by using plane coding.
  • Fig. 5A and Fig. 5B provide a kind of plane position schematic diagram.
  • Fig. 5A shows a kind of low plane position schematic diagram in the Z-axis direction
  • Fig. 5B shows a kind of high plane position schematic diagram in the Z-axis direction.
  • (a), (a0), (a1), (a2), (a3) here all belong to the low plane position in the Z-axis direction.
  • the four subnodes occupied in the current node are located at the high plane position of the current node in the Z-axis direction, so it can be considered that the current node belongs to a Z plane and is a high plane in the Z-axis direction.
  • FIG. 6 provides a schematic diagram of the node coding order, that is, the node coding is performed in the order of 0, 1, 2, 3, 4, 5, 6, and 7 as shown in FIG. 6.
  • the octree coding method is used for (a) in FIG. 5A, the placeholder information of the current node is represented as: 11001100.
  • the plane coding method is used, first, an identifier needs to be encoded to indicate that the current node is a plane in the Z-axis direction.
  • the plane position of the current node needs to be represented; secondly, only the placeholder information of the low plane node in the Z-axis direction needs to be encoded (that is, the placeholder information of the four subnodes 0, 2, 4, and 6). Therefore, based on the plane coding method, only 6 bits need to be encoded to encode the current node, which can reduce the representation of 2 bits compared with the octree coding of the related art. Based on this analysis, plane coding has a more obvious coding efficiency than octree coding.
  • PlaneMode_ i 0 means that the current node is not a plane in the i-axis direction, and 1 means that the current node is a plane in the i-axis direction. If the current node is a plane in the i-axis direction, then for PlanePosition_ i : 0 means that the current node is a plane in the i-axis direction, and the plane position is a low plane, and 1 means that the current node is a high plane in the i-axis direction.
  • Prob(i) new (L ⁇ Prob(i)+ ⁇ (coded node))/L+1 (1)
  • L 255; in addition, if the coded node is a plane, ⁇ (coded node) is 1; otherwise, ⁇ (coded node) is 0.
  • local_node_density new local_node_density+4*numSiblings (2)
  • FIG8 shows a schematic diagram of the sibling nodes of the current node. As shown in FIG8, the current node is a node filled with slashes, and the nodes filled with grids are sibling nodes, then the number of sibling nodes of the current node is 5 (including the current node itself).
  • planarEligibleK OctreeDepth if (pointCount-numPointCountRecon) is less than nodeCount ⁇ 1.3, then planarEligibleK OctreeDepth is true; if (pointCount-numPointCountRecon) is not less than nodeCount ⁇ 1.3, then planarEligibleKOctreeDepth is false. In this way, when planarEligibleKOctreeDepth is true, all nodes in the current layer are plane-encoded; otherwise, all nodes in the current layer are not plane-encoded, and only octree coding is used.
  • Figure 9 shows a schematic diagram of the intersection of a laser radar and a node.
  • a node filled with a grid is simultaneously passed through by two laser beams (Laser), so the current node is not a plane in the vertical direction of the Z axis;
  • a node filled with a slash is small enough that it cannot be passed through by two lasers at the same time, so the node filled with a slash may be a plane in the vertical direction of the Z axis.
  • the plane identification information and the plane position information may be predictively coded.
  • the predictive encoding of the plane position information may include:
  • the plane position information is divided into three elements: predicted as a low plane, predicted as a high plane, and unpredictable;
  • the spatial distance after determining the spatial distance between the node at the same division depth and the same coordinates as the current node and the current node, if the spatial distance is less than a preset distance threshold, then the spatial distance can be determined to be "near”; or, if the spatial distance is greater than the preset distance threshold, then the spatial distance can be determined to be "far”.
  • FIG10 shows a schematic diagram of neighborhood nodes at the same division depth and the same coordinates.
  • the bold large cube represents the parent node (Parent node), the small cube filled with a grid inside it represents the current node (Current node), and the intersection position (Vertex position) of the current node is shown;
  • the small cube filled with white represents the neighborhood nodes at the same division depth and the same coordinates, and the distance between the current node and the neighborhood node is the spatial distance, which can be judged as "near” or "far”; in addition, if the neighborhood node is a plane, then the plane position (Planar position) of the neighborhood node is also required.
  • the current node is a small cube filled with a grid
  • the neighboring node is searched for a small cube filled with white at the same octree partition depth level and the same vertical coordinate, and the distance between the two nodes is judged as "near" and "far", and the plane position of the reference node is referenced.
  • FIG11 shows a schematic diagram of a current node being located at a low plane position of a parent node.
  • (a), (b), and (c) show three examples of the current node being located at a low plane position of a parent node.
  • the specific description is as follows:
  • FIG12 shows a schematic diagram of a current node being located at a high plane position of a parent node.
  • (a), (b), and (c) show three examples of the current node being located at a high plane position of a parent node.
  • the specific description is as follows:
  • Figure 13 shows a schematic diagram of predictive encoding of the laser radar point cloud plane position information.
  • the laser radar emission angle is ⁇ bottom
  • it can be mapped to the bottom plane (Bottom virtual plane)
  • the laser radar emission angle is ⁇ top
  • it can be mapped to the top plane (Top virtual plane).
  • the plane position of the current node is predicted by using the laser radar acquisition parameters, and the position of the current node intersecting with the laser ray is used to quantify the position into multiple intervals, which is finally used as the context information of the plane position of the current node.
  • the specific calculation process is as follows: Assuming that the coordinates of the laser radar are (x Lidar , y Lidar , z Lidar ), and the geometric coordinates of the current node are (x, y, z), then first calculate the vertical tangent value tan ⁇ of the current node relative to the laser radar, and the calculation formula is as follows:
  • each Laser has a certain offset angle relative to the LiDAR, it is also necessary to calculate the relative tangent value tan ⁇ corr,L of the current node relative to the Laser.
  • the specific calculation is as follows:
  • the relative tangent value tan ⁇ corr,L of the current node is used to predict the plane position of the current node. Specifically, assuming that the tangent value of the lower boundary of the current node is tan( ⁇ bottom ), and the tangent value of the upper boundary is tan( ⁇ top ), the plane position is quantized into 4 quantization intervals according to tan ⁇ corr,L , that is, the context information of the plane position is determined.
  • the octree-based geometric information coding mode only has an efficient compression rate for points with correlation in space.
  • the use of the direct coding model (DCM) can greatly reduce the complexity.
  • DCM direct coding model
  • the use of DCM is not represented by flag information, but is inferred from the parent node and neighbor information of the current node. There are three ways to determine whether the current node is eligible for DCM encoding, as follows:
  • the current node has no sibling child nodes, that is, the parent node of the current node has only one child node, and the parent node of the parent node of the current node has only two occupied child nodes, that is, the current node has at most one neighbor node.
  • the parent node of the current node has only one child node, the current node.
  • the six neighbor nodes that share a face with the current node are also empty nodes.
  • FIG14 provides a schematic diagram of IDCM coding. If the current node does not have the DCM coding qualification, it will be divided into octrees. If it has the DCM coding qualification, the number of points contained in the node will be further determined. When the number of points is less than a threshold value (for example, 2), the node will be DCM-encoded, otherwise the octree division will continue.
  • a threshold value for example, 2
  • IDCM_flag the current node is encoded using DCM, otherwise octree coding is still used.
  • the DCM coding mode of the current node needs to be encoded.
  • DCM modes There are currently two DCM modes, namely: (a) only one point exists (or multiple points, but they are repeated points); (b) contains two points.
  • the geometric information of each point needs to be encoded. Assuming that the side length of the node is 2d , d bits are required to encode each component of the geometric coordinates of the node, and the bit information is directly encoded into the bit stream. It should be noted here that when encoding the lidar point cloud, the three-dimensional coordinate information can be predictively encoded by using the lidar acquisition parameters, thereby further improving the encoding efficiency of the geometric information.
  • the current node does not meet the requirements of the DCM node, it will exit directly (that is, the number of points is greater than 2 points and it is not a duplicate point).
  • the second point of the current node is a repeated point, and then it is encoded whether the number of repeated points of the current node is greater than 1. When the number of repeated points is greater than 1, it is necessary to perform exponential Golomb decoding on the remaining number of repeated points.
  • the coordinate information of the points contained in the current node is encoded.
  • the following will introduce the lidar point cloud and the human eye point cloud in detail.
  • the axis with the smaller node coordinate geometry position will be used as the priority coded axis dirextAxis, and then the geometry information of the priority coded axis dirextAxis will be encoded as follows. Assume that the bit depth of the coded geometry corresponding to the priority coded axis is nodeSizeLog2, and assume that the coordinates of the two points are pointPos[0] and pointPos[1].
  • the specific encoding process is as follows:
  • the priority coded coordinate axis dirextAxis geometry information is first encoded as follows, assuming that the priority coded axis corresponds to the coded geometry bit depth of nodeSizeLog2, and assuming that the coordinates of the two points are pointPos[0] and pointPos[1].
  • the specific encoding process is as follows:
  • the geometric coordinate information of the current node can be predicted, so as to further improve the efficiency of the geometric information encoding of the point cloud.
  • the geometric information nodePos of the current node is first used to obtain a directly encoded main axis direction, and then the geometric information of the encoded direction is used to predict the geometric information of another dimension.
  • the axis direction of the direct encoding is directAxis, and the bit depth of the direct encoding is nodeSizeLog2.
  • the encoding is as follows:
  • FIG15 provides a schematic diagram of coordinate transformation of a rotating laser radar to obtain a point cloud.
  • the (x, y, z) coordinates of each node can be converted to (R, i).
  • the laser scanner can perform laser scanning at a preset angle, and different ⁇ (i) can be obtained under different values of i.
  • ⁇ (1) can be obtained, and the corresponding scanning angle is -15°; when i is equal to 2, ⁇ (2) can be obtained, and the corresponding scanning angle is -13°; when i is equal to 10, ⁇ (10) can be obtained, and the corresponding scanning angle is +13°; when i is equal to 9, ⁇ (19) can be obtained, and the corresponding scanning angle is +15°.
  • the LaserIdx corresponding to the current point i.e., the pointLaserIdx number in Figure 15, will be calculated first, and the LaserIdx of the current node, i.e., nodeLaserIdx, will be calculated; secondly, the LaserIdx of the node, i.e., nodeLaserIdx, will be used to predictively encode the LaserIdx of the point, i.e., pointLaserIdx, where the calculation method of the LaserIdx of the node or point is as follows.
  • the LaserIdx of the current node is first used to predict the pointLaserIdx of the point. After the LaserIdx of the current point is encoded, the three-dimensional geometric information of the current point is predicted and encoded using the acquisition parameters of the laser radar.
  • FIG16 shows a schematic diagram of predictive coding in the X-axis or Y-axis direction.
  • a box filled with a grid represents a current node
  • a box filled with a slash represents an already coded node.
  • the LaserIdx corresponding to the current node is first used to obtain the corresponding predicted value of the horizontal azimuth, that is, Secondly, the node geometry information corresponding to the current point is used to obtain the horizontal azimuth angle corresponding to the node Assuming the geometric coordinates of the node are nodePos, the horizontal azimuth
  • the calculation method between the node geometry information is as follows:
  • Figure 17A shows a schematic diagram of predicting the angle of the Y plane through the horizontal azimuth angle
  • Figure 17B shows a schematic diagram of predicting the angle of the X plane through the horizontal azimuth angle.
  • the predicted value of the horizontal azimuth angle corresponding to the current point The calculation is as follows:
  • FIG18 shows another schematic diagram of predictive coding in the X-axis or Y-axis direction.
  • the portion filled with a grid represents the low plane
  • the portion filled with dots represents the high plane.
  • Indicates the horizontal azimuth of the low plane of the current node Indicates the horizontal azimuth of the high plane of the current node, Indicates the predicted horizontal azimuth angle corresponding to the current node.
  • int context (angLel ⁇ 0&&angLeR ⁇ 0)
  • the LaserIdx corresponding to the current point will be used to predict the Z-axis direction of the current point. That is, the depth information radius of the radar coordinate system is calculated by using the x and y information of the current point. Then, the tangent value of the current point and the vertical offset are obtained by using the laser LaserIdx of the current point, and the predicted value of the Z-axis direction of the current point, namely Z_pred, can be obtained.
  • Z_pred is used to perform predictive coding on the geometric information of the current point in the Z-axis direction to obtain the prediction residual Z_res, and finally Z_res is encoded.
  • G-PCC currently introduces a plane coding mode. In the process of geometric division, it will determine whether the child nodes of the current node are in the same plane. If the child nodes of the current node meet the conditions of the same plane, the child nodes of the current node will be represented by the plane.
  • the decoding end follows the order of breadth-first traversal. Before decoding the placeholder information of each node, it will first use the reconstructed geometric information to determine whether the current node is to be plane decoded or IDCM decoded. If the current node meets the conditions for plane decoding, the plane identification and plane position information of the current node will be decoded first, and then the placeholder information of the current node will be decoded based on the plane information; if the current node meets the conditions for IDCM decoding, it will first decode whether the current node is a true IDCM node.
  • IDCM decoding If it is a true IDCM decoding, it will continue to parse the DCM decoding mode of the current node, and then the number of points in the current DCM node can be obtained, and finally the geometric information of each point will be decoded.
  • the placeholder information of the current node will be decoded.
  • the prior information is first used to determine whether the node starts IDCM. That is, the starting conditions of IDCM are as follows:
  • the current node has no sibling child nodes, that is, the parent node of the current node has only one child node, and the parent node of the parent node of the current node has only two occupied child nodes, that is, the current node has at most one neighbor node.
  • the parent node of the current node has only one child node, the current node.
  • the six neighbor nodes that share a face with the current node are also empty nodes.
  • a node meets the conditions for DCM coding, first decode whether the current node is a real DCM node, that is, IDCM_flag; when IDCM_flag is true, the current node adopts DCM coding, otherwise it still adopts octree coding.
  • numPonts of the current node obtained by decoding is less than or equal to 1, continue decoding to see if the second point is a repeated point; if the second point is not a repeated point, it can be implicitly inferred that the second type that satisfies the DCM mode contains only one point; if the second point obtained by decoding is a repeated point, it can be inferred that the third type that satisfies the DCM mode contains multiple points, but they are all repeated points, then continue decoding to see if the number of repeated points is greater than 1 (entropy decoding), and if it is greater than 1, continue decoding the number of remaining repeated points (decoding using exponential Columbus).
  • the current node does not meet the requirements of the DCM node, it will exit directly (that is, the number of points is greater than 2 points and it is not a duplicate point).
  • the coordinate information of the points contained in the current node is decoded.
  • the following will introduce the lidar point cloud and the human eye point cloud in detail.
  • the axis with the smaller node coordinate geometry position will be used as the priority decoding axis dirextAxis, and then the priority decoding axis dirextAxis geometry information will be decoded first in the following way.
  • the geometry bit depth to be decoded corresponding to the priority decoding axis is nodeSizeLog2
  • the coordinates of the two points are pointPos[0] and pointPos[1] respectively.
  • the specific encoding process is as follows:
  • the priority encoded coordinate axis dirextAxis geometry information is first decoded as follows, assuming that the priority decoded axis corresponds to the code geometry bit depth of nodeSizeLog2, and assuming that the coordinates of the two points are pointPos[0] and pointPos[1].
  • the specific encoding process is as follows:
  • the LaserIdx of the current node i.e., nodeLaserIdx
  • the LaserIdx of the node i.e., nodeLaserIdx
  • the calculation method of the LaserIdx of the node or point is the same as that of the encoder.
  • the LaserIdx of the current point and the predicted residual information of the LaserIdx of the node are decoded to obtain ResLaserIdx.
  • the three-dimensional geometric information of the current point is predicted and decoded using the acquisition parameters of the laser radar.
  • the specific algorithm is as follows:
  • the node geometry information corresponding to the current point is used to obtain the horizontal azimuth angle corresponding to the node Assuming the geometric coordinates of the node are nodePos, the horizontal azimuth
  • the calculation method between the node geometry information is as follows:
  • int context (angLel ⁇ 0&&angLeR ⁇ 0)
  • the Z-axis direction of the current point will be predicted and decoded using the LaserIdx corresponding to the current point, that is, the depth information radius of the radar coordinate system is calculated by using the x and y information of the current point, and then the tangent value of the current point and the vertical offset are obtained using the laser LaserIdx of the current point, so the predicted value of the Z-axis direction of the current point, namely Z_pred, can be obtained.
  • the decoded Z_res and Z_pred are used to reconstruct and restore the geometric information of the current point in the Z-axis direction.
  • geometric division For the geometric information coding based on triangle soup (trisoup), in the geometric information coding framework based on trisoup, geometric division must also be performed first, but different from the geometric information coding based on binary tree/quadtree/octree, this method does not
  • the point cloud needs to be divided into unit cubes with a side length of 1 ⁇ 1 ⁇ 1 step by step, and the division stops when the side length of the sub-block is W.
  • the surface and the twelve edges of the block are used to obtain at most twelve intersection points (vertex).
  • the vertex coordinates of each block are encoded in turn to generate a binary code stream.
  • the Predictive geometry coding includes: first, sorting the input point cloud.
  • the currently used sorting methods include unordered, Morton order, azimuth order, and radial distance order.
  • the prediction tree structure is established by using two different methods, including: KD-Tree (high-latency slow mode) and low-latency fast mode (using laser radar calibration information).
  • KD-Tree high-latency slow mode
  • low-latency fast mode using laser radar calibration information.
  • each node in the prediction tree is traversed, and the geometric position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometric prediction residual is quantized using the quantization parameter.
  • the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary code stream.
  • the decoding end reconstructs the prediction tree structure by continuously parsing the bit stream, and then obtains the geometric position prediction residual information and quantization parameters of each prediction node through parsing, and dequantizes the prediction residual to recover the reconstructed geometric position information of each node, and finally completes the geometric reconstruction of the decoding end.
  • attribute encoding is mainly performed on color information.
  • the color information is converted from the RGB color space to the YUV color space.
  • the point cloud is recolored using the reconstructed geometric information so that the unencoded attribute information corresponds to the reconstructed geometric information.
  • color information encoding there are two main transformation methods, one is the distance-based lifting transformation that relies on LOD division, and the other is to directly perform RAHT transformation. Both methods will convert color information from the spatial domain to the frequency domain, and obtain high-frequency coefficients and low-frequency coefficients through transformation.
  • the coefficients are quantized and encoded to generate a binary code stream, as shown in Figures 4A and 4B.
  • the Morton code can be used to search for the nearest neighbor.
  • the Morton code corresponding to each point in the point cloud can be obtained from the geometric coordinates of the point.
  • the specific method for calculating the Morton code is described as follows. For each component of the three-dimensional coordinate represented by a d-bit binary number, its three components can be expressed as:
  • the highest bits of x, y, and z are To the lowest position The corresponding binary value.
  • the Morton code M is x, y, z, starting from the highest bit, arranged in sequence To the lowest bit, the calculation formula of M is as follows:
  • Condition 1 The geometric position is limitedly lossy and the attributes are lossy;
  • Condition 3 The geometric position is lossless, and the attributes are limitedly lossy
  • Condition 4 The geometric position and attributes are lossless.
  • the general test sequences include four categories: Cat1A, Cat1B, Cat3-fused, and Cat3-frame.
  • the Cat2-frame point cloud only contains reflectance attribute information
  • the Cat1A and Cat1B point clouds only contain color attribute information
  • the Cat3-fused point cloud contains both color and reflectance attribute information.
  • the bounding box is divided into sub-cubes in sequence, and the non-empty sub-cubes (containing points in the point cloud) are divided again until the leaf node obtained by division is a 1 ⁇ 1 ⁇ 1 unit cube.
  • the number of points contained in the leaf node needs to be encoded, and finally the encoding of the geometric octree is completed to generate a binary code stream.
  • the decoding end obtains the placeholder code of each node by continuously parsing in the order of breadth-first traversal, and continuously divides the nodes in turn until a 1 ⁇ 1 ⁇ 1 unit cube is obtained.
  • geometric lossless decoding it is necessary to parse the number of points contained in each leaf node and finally restore the geometrically reconstructed point cloud information.
  • the prediction tree structure is established by using two different methods, including: based on KD-Tree (high-latency slow mode) and using lidar calibration information (low-latency fast mode).
  • lidar calibration information each point can be divided into different Lasers, and the prediction tree structure is established according to different Lasers.
  • each node in the prediction tree is traversed, and the geometric position information of the node is predicted by selecting different prediction modes to obtain the prediction residual, and the geometric prediction residual is quantized using the quantization parameter.
  • the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary code stream.
  • the decoding end reconstructs the prediction tree structure by continuously parsing the bit stream, and then obtains the geometric position prediction residual information and quantization parameters of each prediction node through parsing, and dequantizes the prediction residual to restore the reconstructed geometric position information of each node, and finally completes the geometric reconstruction at the decoding end.
  • the current G-PCC coding framework includes three attribute coding methods: Predicting Transform (PT), Lifting Transform (LT), and Region Adaptive Hierarchical Transform (RAHT).
  • PT Predicting Transform
  • LT Lifting Transform
  • RAHT Region Adaptive Hierarchical Transform
  • the first two predict the point cloud based on the generation order of LOD
  • RAHT adaptively transforms the attribute information from bottom to top based on the construction level of the octree.
  • PT Predicting Transform
  • LT Lifting Transform
  • RAHT Region Adaptive Hierarchical Transform
  • the attribute prediction module of G-PCC adopts a nearest neighbor attribute prediction coding scheme based on a hierarchical (Level-of-details, LoDs) structure.
  • the LOD construction methods include distance-based LOD construction schemes, fixed sampling rate-based LOD construction schemes, and octree-based LOD construction schemes.
  • the point cloud is first Morton sorted before constructing the LOD to ensure that there is a strong attribute correlation between adjacent points.
  • Rl point cloud detail layers
  • the attribute value of each point is linearly weighted predicted by using the attribute reconstruction value of the point in the same layer or higher LOD, where the maximum number of reference prediction neighbors is determined by the encoder high-level syntax elements.
  • the encoding end uses the rate-distortion optimization algorithm to select the weighted prediction by using the attributes of the N nearest neighbor points searched or the attribute of a single nearest neighbor point for prediction, and finally encodes the selected prediction mode and prediction residual.
  • N represents the number of predicted points in the nearest neighbor point set of point i
  • Pi represents the sum of the N nearest neighbor points of point i
  • Dm represents the spatial geometric distance from the nearest neighbor point m to the current point i
  • Attrm represents the attribute value after reconstruction of the nearest neighbor point m
  • Attr i ′ represents the attribute prediction value of the current point i
  • the number of points N is a preset value.
  • a switch is introduced in the encoder high-level syntax element to control whether to introduce LOD layer intra prediction. If it is turned on, LOD layer intra prediction is enabled, and points in the same LOD layer can be used for prediction. It should be noted that when the number of LOD layers is 1, LOD layer intra prediction is always used.
  • FIG21 is a schematic diagram of a visualization result of the LOD generation process. As shown in FIG21, a subjective example of the distance-based LOD generation process is provided. Specifically (from left to right): the points in the first layer represent the outer contour of the point cloud; as the number of detail layers increases, the point cloud detail description becomes clearer.
  • Figure 22 is a schematic diagram of the encoding process of attribute prediction.
  • attribute prediction for the specific process of G-PCC attribute prediction, for the original point cloud, first search for the three neighboring points of the Kth point, and then perform attribute prediction; calculate the difference between the attribute prediction value of the Kth point and the original attribute value of the Kth point to obtain the prediction residual of the Kth point; then perform quantization and arithmetic coding to finally generate the attribute bit rate.
  • the LOD After the LOD is constructed, according to the generation order of LOD, first find the three nearest neighbor points of the current point to be encoded from the encoded data points. The attribute reconstruction values of these three nearest neighbor points are used as candidate prediction values of the current point to be encoded; then, the optimal prediction value is selected from them according to the rate-distortion optimization (RDO).
  • RDO rate-distortion optimization
  • the prediction variable index of the attribute value of the nearest neighbor point P4 is set to 1; the attribute prediction variable indexes of the second nearest neighbor point P5 and the third nearest neighbor point P0 are set to 2 and 3 respectively; the prediction variable index of the weighted average of points P0, P5 and P4 is set to 0, as shown in Table 1; finally, using RDO Select the best predictor variable.
  • the formula for weighted average is as follows:
  • x i , y i , zi are the geometric position coordinates of the current point i
  • x ij , y ij , z ij are the geometric coordinates of the neighboring point j.
  • Table 1 provides an example of a sample of candidate prediction items for an attribute encoding.
  • the attribute prediction value of the current point i is obtained through the above prediction (k is the total number of points in the point cloud).
  • (a i ) i ⁇ 0...k-1 be the original attribute value of the current point, then the attribute residual (r i ) i ⁇ 0...k-1 is recorded as:
  • the prediction residuals are further quantified:
  • Qi represents the quantized attribute residual of the current point i
  • Qs is the quantization step (Quantization step, Qs), which can be calculated by the quantization parameter QP (Quantization Parameter, QP) specified by CTC.
  • the purpose of reconstruction at the encoding end is to predict subsequent points. Before reconstructing the attribute value, the residual must be dequantized. is the residual after inverse quantization:
  • intra-frame nearest neighbor search When performing attribute nearest neighbor search based on LOD division, there are currently two major types of algorithms: intra-frame nearest neighbor search and inter-frame nearest neighbor search.
  • inter-frame nearest neighbor search algorithm is as follows, and the intra-frame nearest neighbor search can be divided into two algorithms: inter-layer nearest neighbor search and intra-layer nearest neighbor search.
  • the nearest neighbor search within a frame is divided into two algorithms: the inter-layer nearest neighbor search and the intra-layer nearest neighbor search. After LOD division, it is similar to a pyramid structure, as shown in Figure 23.
  • FIG24 is a pyramid structure for inter-layer nearest neighbor search.
  • LOD0, LOD1 and LOD2 use the points in LOD0 to predict the attributes of the points in the next layer of LOD in the nearest neighbor search between layers
  • the entire LOD division process there are three sets O(k), L(k) and I(k). Among them, k is the index of the LOD layer during LOD division, I(k) is the input point set during the current LOD layer division, and after LOD division, O(k) set and L(k) set are obtained. The O(k) set stores the sampling point set, and L(k) is the point set in the current LOD layer. That is, the entire LOD division process is as follows:
  • O(k), L(k) and I(k) store the Morton code index corresponding to the point.
  • the neighbor search is performed by using the parent block (Block B) corresponding to point P, as shown in Figure 26, and the points in the neighbor blocks that are coplanar and colinear with the current parent block are searched for attribute prediction.
  • FIG. 27A shows a schematic diagram of a coplanar spatial relationship, where there are 6 spatial blocks that have a relationship with the current parent block.
  • FIG. 27B shows a schematic diagram of a coplanar and colinear spatial relationship, where there are 18 spatial blocks that have a relationship with the current parent block.
  • FIG. 27C shows a schematic diagram of a coplanar, colinear and co-point spatial relationship, where there are 26 spatial blocks that have a relationship with the current parent block.
  • the coordinates of the current point are used to obtain the corresponding spatial block.
  • the nearest neighbor search is performed in the previously encoded LOD layer to find the spatial blocks that are coplanar, colinear, and co-point with the current block to obtain the N nearest neighbors of the current point.
  • the N nearest neighbors of the current point After searching for coplanar, colinear, and co-point nearest neighbors, if the N nearest neighbors of the current point are still not found, the N nearest neighbors of the current point will be found based on the fast search algorithm.
  • the specific algorithm is as follows:
  • the geometric coordinates of the current point to be encoded are first used to obtain the Morton code corresponding to the current point. Secondly, based on the Morton code of the current point, the first reference point (j) that is larger than the Morton code of the current point is found in the reference frame. Then, the nearest neighbor search is performed in the range of [j-searchRange, j+searchRange].
  • FIG29 shows a schematic diagram of the LOD structure of the nearest neighbor search within an attribute layer.
  • the nearest neighbor point of the current point P6 can be P4.
  • the nearest neighbor search will be performed in the same layer LOD and the set of encoded points in the same layer to obtain the N nearest neighbors of the current point (inter-layer nearest neighbor search is also performed).
  • the nearest neighbor search is performed based on the fast search algorithm.
  • the specific algorithm is shown in Figure 30.
  • the current point is represented by a grid.
  • the nearest neighbor search is performed in [i+1, i+searchRange].
  • the specific nearest neighbor search algorithm is consistent with the inter-frame block-based fast search algorithm and will not be described in detail here.
  • Figure 28 is a schematic diagram of attribute inter-frame prediction.
  • attribute inter-frame prediction when performing attribute inter-frame prediction, firstly, the geometric coordinates of the current point to be encoded are used to obtain the Morton code corresponding to the current point, and then the first reference point (j) with a value greater than the Morton code of the current point is found in the reference frame based on the Morton code of the current point, and then the nearest neighbor search is performed within the range of [j-searchRange, j+searchRange].
  • the specific division algorithm is as follows:
  • the reference range in the prediction frame of the current point is [j-searchRange, j+searchRange], use j-searchRange to calculate the starting index of the third layer, and use j+searchRange to calculate the ending index of the third layer; secondly, first determine whether some blocks in the second layer need to be searched for the nearest neighbor in the blocks of the third layer, and then go to the second layer, and determine whether a search is needed for each block in the first layer. If some blocks in the first layer need to be searched for the nearest neighbor, then some midpoints of some blocks in the first layer will be judged point by point to update the nearest neighbors.
  • the index of the first layer block is obtained based on the index of the second layer block based on the same algorithm.
  • MinPos represents the minimum value of the block
  • maxPos represents the maximum value of the block.
  • the coordinates of the point to be encoded are (x, y, z), and the current block is represented by (minPos, maxPos), where minPos is the minimum value of the bounding box in three dimensions, and maxPos is the maximum value of the bounding box in three dimensions.
  • Figure 32 is a schematic diagram of the encoding process of a lifting transformation.
  • the lifting transformation also predicts the attributes of the point cloud based on LOD.
  • the difference from the prediction transformation is that the lifting transformation first divides the LOD into high and low layers, predicts in the reverse order of the LOD generation layer, and introduces an update operator in the prediction process to update the quantization weights of the midpoints of the low-level LOD to improve the accuracy of the prediction. This is because the attribute values of the midpoints of the low-level LOD are frequently used to predict the attribute values of the midpoints of the high-level LOD, and the points in the low-level LOD should have greater influence.
  • Step 1 Segmentation process.
  • Step 2 Prediction process.
  • Step 3 Update Process.
  • the transformation scheme based on lifting wavelet transform introduces quantization weights and updates the prediction residual according to the prediction residual D(N) and the distance between the prediction point and the adjacent points, and finally uses the quantization weights in the transformation process to adaptively quantize the prediction residual.
  • the quantization weight value of each point can be determined by geometric reconstruction at the decoding end, so the quantization weight should not be encoded.
  • Regional Adaptive Hierarchical Transform is a Haar wavelet transform that can transform point cloud attribute information from the spatial domain to the frequency domain, further reducing the correlation between point cloud attributes. Its main idea is to transform the nodes in each layer from the three dimensions of X, Y, and Z in a bottom-up manner according to the octree structure (as shown in Figure 34), and iterate until the root node of the octree. As shown in Figure 33, its basic idea is to perform wavelet transform based on the hierarchical structure of the octree, associate attribute information with the octree nodes, and recursively transform the attributes of the occupied nodes in the same parent node in a bottom-up manner.
  • RAHT Regional Adaptive Hierarchical Transform
  • the nodes are transformed from the three dimensions of X, Y, and Z until they are transformed to the root node of the octree.
  • the low-pass/low-frequency (DC) coefficients obtained after the transformation of the nodes in the same layer are passed to the nodes in the next layer for further transformation, and all high-pass/high-frequency (AC) coefficients can be encoded by the arithmetic encoder.
  • the DC coefficient (direct current component) of the nodes in the same layer after transformation will be transferred to the previous layer for further transformation, and the AC coefficient (alternating current component) after transformation in each layer will be quantized and encoded.
  • the main transformation process will be introduced below.
  • FIG35A is a schematic diagram of a RAHT forward transformation process
  • FIG35B is a schematic diagram of a RAHT inverse transformation process.
  • g′ L,2x,y,z and g′ L,2x+1,y,z are two attribute DC coefficients that are neighboring points in the L layer.
  • the information of the L-1 layer is the AC coefficient f′ L-1,x,y,z and the DC coefficient g′ L-1,x,y,z ; then, f′ L-1,x,y,z will no longer be transformed and will be directly quantized and encoded, and g′ L-1,x,y,z will continue to look for neighbors for transformation.
  • T w0,w1 is the transformation matrix:
  • the transformation matrix will be updated as the weights corresponding to each point change adaptively.
  • the above process will be iteratively updated according to the partition structure of the octree until the root node of the octree.
  • the attribute is determined by the attribute parameter set (APS) syntax element to determine which inter-frame prediction coding scheme to use for inter-frame prediction of the attribute, and a syntax element treeDepth is used to determine the number of layers to start the inter-frame prediction coding.
  • APS attribute parameter set
  • treeDepth is used to determine the number of layers to start the inter-frame prediction coding.
  • inter-frame coding is often only started in the upper layer of RAHT coding.
  • such a coding scheme does not fully and effectively utilize the distribution of AC coefficients in different RAHT layers, resulting in low coding efficiency of attribute information.
  • an embodiment of the present application provides a decoding method, which parses the bitstream and determines the first syntax identification information when it is determined that the nodes of the current layer allow attribute prediction; when the first syntax identification information indicates that the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode, parses the bitstream and determines the target decoding mode of the current layer; and performs attribute decoding on the nodes in the current layer according to the target decoding mode to determine the attribute reconstruction values of the nodes in the current layer.
  • An embodiment of the present application also provides a coding method, which determines the target coding mode of the current layer and determines the first grammar identification information when it is determined that the nodes of the current layer allow attribute prediction and the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode; the first grammar identification information is used to indicate whether the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode; attributes of the nodes in the current layer are encoded according to the target coding mode, and attribute reconstruction values of the nodes in the current layer are determined.
  • the encoding end when performing attribute encoding for each layer, can adaptively select the target coding mode of each slice, and pass the target coding mode to the decoding end, so that the decoding end uses the parsed target decoding mode to reconstruct the attributes of the point cloud, thereby improving the coding efficiency of the point cloud attributes, and then improving the encoding and decoding performance of the point cloud.
  • FIG36 a schematic flow chart of a decoding method provided by an embodiment of the present application is shown. As shown in FIG36 , the method may include S101 to S103:
  • the decoding method is applied to a point cloud decoder (hereinafter referred to as "decoder").
  • the decoding method may be a point cloud attribute decoding method, and more specifically, may be a method for adaptively selecting inter-frame prediction or intra-frame prediction for decoding based on point cloud attribute RAHT transform prediction.
  • the corresponding attribute decoding mode is mainly introduced for each layer in the current sequence in the attribute block header information parameter set (Attribute Brick Header, ABH), and the corresponding target decoding mode can be adaptively selected for each layer (Layer), thereby improving the decoding efficiency of point cloud attributes.
  • the current layer may be one of the layers in the current video frame.
  • a video frame may be understood as an image.
  • a current frame may be understood as a current image
  • a reference frame may be understood as a reference image.
  • the current layer includes at least one node.
  • the current layer may be referred to as the current attribute decoding layer, the current decoding layer, the current slice, etc.
  • the embodiment of the present application does not impose any limitation on this.
  • the current layer is a decoding layer obtained by upsampling along the first direction, the second direction and the third direction, wherein the first direction is the z-axis direction, the second direction is the y-axis direction, and the third direction is the x-axis direction.
  • the embodiment of the present application does not limit the order of the first direction, the second direction, and the third direction.
  • it can be the second direction, the first direction, and the third direction, or it can be the third direction, the second direction, and the first direction.
  • the current layer is not limited to a decoding layer obtained by upsampling once along the first direction, the second direction, and the third direction.
  • the current layer may also be multiple decoding layers obtained by upsampling once along the first direction, the second direction, and the third direction.
  • the current layer may also be a layer composed of at least one node in a decoding layer. The embodiment of the present application does not impose any limitation on this.
  • the first syntax identification information is used to indicate that the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode.
  • the implementation of parsing the code stream and determining the value of the first syntax identification information may include:
  • the current coefficient group allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode; the current coefficient group includes at least one layer, and the current layer is one of the at least one layer;
  • the value of the first syntax identification information is the second value, it is determined that the current coefficient group does not allow adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode.
  • the current coefficient group includes at least one layer, and the current layer is one of the at least one layer.
  • the implementation of parsing the code stream and determining the value of the first syntax identification information may include:
  • the value of the first syntax identification information is the first value, it is determined that the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode;
  • the value of the first syntax identification information is the second value, it is determined that the current layer does not allow adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode.
  • the first value is different from the second value, and the first value and the second value can be in parameter form or in digital form.
  • the first syntax identification information can be a parameter written in the profile or a flag value, which is not specifically limited here.
  • the first value can be set to 1 and the second value can be set to 0; or, the first value can be set to 0 and the second value can be set to 1; or, the first value can be set to true and the second value can be set to false; or, the first value can be set to false and the second value can be set to true; but this is not specifically limited here.
  • the first grammar identification information acts as a switch, that is, when the first grammar identification information is a first value (such as 1 or true), it indicates that the decoding algorithm described in the embodiment of the present application is started, that is, the decoding algorithm described in the embodiment of the present application is executed; when the first grammar identification information is a second value (such as 0 or false), it indicates that the decoding algorithm described in the embodiment of the present application is not started, that is, the decoding algorithm described in the embodiment of the present application is not executed.
  • a first value such as 1 or true
  • different layers can adaptively select intra-frame prediction and/or inter-frame prediction to perform attribute decoding on the node, which can fully consider the distribution of AC coefficients of different layers, thereby improving the decoding efficiency of RAHT.
  • determining whether the node of the current layer is allowed to perform attribute prediction may include the following two methods:
  • Method 1 Determine, based on the sixth grammar identification information, whether the node in the current layer is allowed to perform attribute prediction;
  • the implementation of method 1 may include:
  • the bitstream Parse the bitstream to determine sixth syntax identification information; the sixth syntax identification information is used to indicate whether a node in the current layer is allowed to perform attribute prediction;
  • the value of the sixth grammar identification information is the first value, it is determined that the node of the current layer allows attribute prediction
  • the value of the first syntax identification information is the second value, it is determined that the node of the current layer is not allowed to perform attribute prediction.
  • the decoding end adopts method 1 to determine whether the node of the current layer is allowed to perform attribute prediction, it is only necessary to parse the code stream and make a judgment based on the value of the sixth grammar identification information. In this way, the decoding end can avoid the need to repeat the judgment process, simplify the decoding process, and thus improve the decoding efficiency.
  • Method 2 According to the number of adjacent nodes in the current layer, determine whether the nodes in the current layer are allowed to perform attribute prediction.
  • the implementation of method 2 may include:
  • the decoding end and the encoding end adopt the same process to determine In this way, the encoder does not need to transmit codewords indicating whether the nodes in the current layer are allowed to perform attribute prediction to the decoder, and the decoder does not need to parse the corresponding codewords, which can also improve the decoding efficiency to a certain extent.
  • the decoder determines the target decoding mode corresponding to the current layer by parsing the bitstream.
  • the target decoding mode can be expressed as attr_code_mode[i]; where i is the index value of the current layer.
  • index value i is assigned only when the current layer satisfies the three conditions of allowing attribute prediction, allowing inter-frame prediction, and allowing intra-frame prediction.
  • the decoder directly skips the current layer, directly performs attribute decoding of the next layer, and assigns index 2 to the next layer; if the current layer meets the three conditions of allowing attribute prediction, allowing inter-frame prediction, and allowing intra-frame prediction, the decoder performs attribute decoding on the current layer, adds 1 to the index value (i++), obtains the updated index value (3), and passes the index value 3 to the next layer.
  • parsing the bitstream in S102 to determine the implementation of the target decoding mode of the current layer may include S1021 to S1023:
  • an attribute block header information parameter set may include a target decoding mode corresponding to at least one layer.
  • S1022 Determine second syntax identification information from the attribute block header information parameter set.
  • the decoder determines the second syntax identification information corresponding to the current layer from the attribute block header information parameter set, wherein the second syntax identification information is used to indicate the target decoding mode of the current layer.
  • the decoder determines the target decoding mode of the current layer according to the value of the second syntax identification information corresponding to the current layer.
  • the value of the second syntax identification information may be in parameter form or in digital form, and the embodiment of the present application does not impose any limitation on this.
  • the target decoding mode includes a region adaptive hierarchical intra-frame transform mode, a region adaptive hierarchical inter-frame transform mode and a region adaptive hierarchical combined transform mode;
  • the regional adaptive hierarchical intra-frame transform mode representation adopts the intra-frame prediction mode to perform attribute prediction transform decoding on the nodes of the current layer;
  • the regional adaptive hierarchical inter-frame transform mode representation adopts the inter-frame prediction mode to perform attribute prediction transform decoding on the nodes of the current layer;
  • the regional adaptive hierarchical combined transform mode representation adopts the intra-frame prediction mode combined with the inter-frame prediction mode to perform attribute prediction transform decoding on the nodes of the current layer.
  • the regional adaptive hierarchical inter-frame transform mode includes a first regional adaptive hierarchical inter-frame transform mode and a second regional adaptive hierarchical inter-frame transform mode;
  • the regional adaptive hierarchical combined transform mode includes a first regional adaptive hierarchical combined transform mode, a second regional adaptive hierarchical combined transform mode and a third regional adaptive hierarchical combined transform mode;
  • the first region adaptive hierarchical inter-frame transform mode represents the method of using the geometric information of the node to determine the co-located prediction node to perform attribute prediction transform decoding on the node of the current layer;
  • the second region adaptive hierarchical combination transform mode characterization uses the reference frame cache to determine the same-position prediction node to perform attribute prediction transform decoding on the nodes of the current layer;
  • the first region adaptive hierarchical combined transform mode characterization adopts the combined region adaptive hierarchical intra-frame transform mode and the first region adaptive hierarchical inter-frame transform mode to perform attribute prediction transform decoding on the nodes of the current layer;
  • the second region adaptive hierarchical combined transform mode characterization adopts the combined region adaptive hierarchical intra-frame transform mode and the second region adaptive hierarchical inter-frame transform mode to perform attribute prediction transform decoding on the nodes of the current layer;
  • the third region adaptive hierarchical combined transform mode characterization uses a combined region adaptive hierarchical intra-frame transform mode, a first region adaptive hierarchical inter-frame transform mode, and a second region adaptive hierarchical inter-frame transform mode to perform attribute prediction transform decoding on nodes of the current layer.
  • prediction can be performed based on the RAHT transform codec.
  • the RAHT attribute transform is based on the order of the octree hierarchy, which is continuously transformed at the voxel level. The transformation is performed until the root node is obtained, thereby completing the hierarchical transformation encoding and decoding of the entire attribute.
  • the attribute prediction transformation encoding and decoding is also performed based on the hierarchical order of the octree, but the transformation is continuously performed from the root node until the voxel level.
  • each RAHT attribute transformation process the attribute prediction transformation encoding and decoding is performed based on a 2 ⁇ 2 ⁇ 2 block.
  • the details are shown in Figure 37.
  • the grid filling block is the current block to be encoded and decoded
  • the diagonal filling block is some neighboring blocks that are coplanar and colinear with the current block to be encoded and decoded.
  • the attributes of the current block can be obtained by the attributes of the points in the current block, that is, A node .
  • the attributes of the points in the current block are simply added, and then the attributes of the current block and the number of points in the current block are normalized to obtain the mean value of the attributes of the current block a node .
  • the mean value of the attributes of the current block is used for attribute transformation encoding and decoding. For the specific encoding and decoding process, see Figure 38.
  • RAHT attribute prediction transformation encoding and decoding As shown in Figure 38, the overall process of RAHT attribute prediction transformation encoding and decoding is shown here. Among them, (a) is the current block and some coplanar and colinear neighboring blocks, (b) is the block after normalization, (c) is the block after upsampling, (d) is the attribute of the current block, and (e) is the attribute of the predicted block obtained by linear weighted fitting using the neighborhood attributes of the current block. Finally, the attributes of the two will be transformed respectively to obtain DC and AC coefficients, and the AC coefficient will be predicted and encoded.
  • the predicted attribute of the current block can be obtained by linear fitting as shown in FIG39.
  • FIG39 firstly, 19 neighboring blocks of the current block are obtained, and then the attribute of each sub-block is linearly weighted predicted using the spatial geometric distance between the neighboring block and each sub-block of the current block, and finally the predicted block attribute obtained by linear weighting is transformed.
  • the specific attribute transformation is shown in FIG40.
  • (d) represents the original value of the attribute
  • the corresponding attribute transformation coefficient is as follows:
  • (e) represents the attribute prediction value, and the corresponding attribute transformation coefficient is as follows:
  • the prediction residual By subtracting the original value of the attribute from the predicted value of the attribute, the prediction residual can be obtained as follows:
  • the first region adaptive hierarchical inter-frame transform mode is also called region adaptive hierarchical inter-frame prediction transform coding scheme 1.
  • the process is similar to the intra-frame prediction coding and decoding.
  • the RAHT attribute transform coding and decoding structure is constructed based on the geometric information, that is, the voxel level is continuously transformed until the root node is obtained, thereby completing the hierarchical transform coding and decoding of the entire attribute.
  • the intra-frame coding and decoding structure and the inter-frame attribute coding and decoding structure are constructed, see Figure 41 for details.
  • the geometric information of the current node to be encoded and decoded is used to obtain the co-located predicted node of the node to be encoded and decoded in the reference frame, and then the geometric information and attribute information of the reference node are used to obtain the predicted attribute of the current node to be encoded and decoded.
  • the attribute prediction value of the current node to be encoded and decoded is obtained in the following two different ways:
  • the inter-frame prediction node of the current node is valid: that is, if the same-position node exists, the attribute of the prediction node is directly used as the attribute prediction value of the current node to be encoded and decoded;
  • the inter-frame prediction node of the current node is invalid: that is, the co-located node does not exist, then the attribute prediction value of the adjacent node in the frame is used as the attribute prediction value of the node to be encoded and decoded.
  • the obtained attribute prediction value is used to predict the attribute of the current node to be encoded and decoded, thereby completing the prediction encoding and decoding of the entire attribute.
  • the second region adaptive hierarchical inter-frame transform mode is also called the region adaptive hierarchical inter-frame prediction transform coding and decoding scheme 2.
  • the RAHT attribute transform coding and decoding structure is first constructed based on the geometric information of the current node to be encoded and decoded, that is, the nodes are continuously merged at the voxel level until the root node of the entire RAHT transform tree is obtained. Then, the transformation codec hierarchical structure of the entire attribute is completed.
  • the root node is divided to obtain N child nodes (N is less than or equal to 8) of each node.
  • N is less than or equal to 8
  • the attributes of the N child nodes are firstly orthogonally transformed independently using the RAHT transformation to obtain the DC coefficient (direct current component) and the AC coefficient (alternating current component). Then, the AC coefficients of the N child nodes are predicted for the attributes of the inter-frame according to the following method:
  • the inter-frame prediction node of the current node is valid: that is, if the same-position node exists, the attribute of the prediction node is directly used as the attribute prediction value of the current node to be encoded and decoded;
  • the previous node can find a node with exactly the same position as the current node in the cache of the reference frame: that is, if the same-position node exists, the AC coefficients of the M child nodes contained in the same-position node will be directly used as the AC coefficient attribute prediction values of the N child nodes of the current node.
  • the inter-frame prediction node of the current node is invalid: that is, the co-located node does not exist, then the attribute prediction value of the adjacent node in the frame is used as the attribute prediction value of the node to be encoded and decoded.
  • the implementation of determining the target decoding mode of the current layer according to the second syntax identification information in S1023 may include: determining the target decoding mode of the current layer using the inter-frame prediction mode and/or the intra-frame prediction mode according to the value of the second syntax identification information.
  • determining the implementation of the target decoding mode of the current layer using the inter-frame prediction mode and/or the intra-frame prediction mode may include:
  • the value of the second syntax identification information is the third value, determining that the target decoding mode of the current layer is the region adaptive hierarchical intra transform mode
  • the value of the second syntax identification information is the fourth value, determining that the target decoding mode of the current layer is the first region adaptive hierarchical inter-frame transform mode
  • the target decoding mode of the current layer is the second region adaptive hierarchical inter-frame transform mode
  • the value of the second syntax identification information is the sixth value, determining that the target decoding mode of the current layer is the first region adaptive hierarchical combined transform mode
  • the value of the second syntax identification information is the seventh value, determining that the target decoding mode of the current layer is the second region adaptive hierarchical combined transform mode;
  • the target decoding mode of the current layer is determined to be the third region adaptive hierarchical combined transform mode.
  • the third value, the fourth value, the fifth value, the sixth value, the seventh value and the eighth value are different. It should be noted that the third value, the fourth value, the fifth value, the sixth value, the seventh value and the eighth value can be in parameter form or in digital form. Exemplarily, the third value is 0, the fourth value is 1, the fifth value is 2, the sixth value is 3, the seventh value is 4, and the eighth value is 5. The embodiment of the present application does not impose any restrictions on the setting of the third value, the fourth value, the fifth value, the sixth value, the seventh value and the eighth value.
  • S103 Perform attribute decoding on the nodes in the current layer according to the target decoding mode to determine the attribute reconstruction values of the nodes in the current layer.
  • the decoder can perform attribute decoding on the nodes in the current layer according to the target decoding mode, and then determine the attribute reconstruction values of the nodes in the current layer.
  • the decoder if the target decoding mode is the region adaptive layered intra-frame transform mode, the decoder performs attribute decoding on the nodes in the current layer according to the region adaptive layered intra-frame transform mode, and then determines the attribute reconstruction values of the nodes in the current layer; if the target decoding mode is the first region adaptive layered inter-frame transform mode, the decoder performs attribute decoding on the nodes in the current layer according to the first region adaptive layered inter-frame transform mode, and then determines the attribute reconstruction values of the nodes in the current layer; if the target decoding mode is the second region adaptive layered inter-frame transform mode, the decoder performs attribute decoding on the nodes in the current layer according to the second region adaptive layered inter-frame transform mode, and then determines the attribute reconstruction values of the nodes in the current layer.
  • the target decoding mode is the first region adaptive layered combined transform mode
  • the decoder performs attribute decoding on the nodes in the current layer according to the first region adaptive layered combined transform mode, and then determines the attribute reconstruction value of the nodes in the current layer
  • the target decoding mode is the second region adaptive layered combined transform mode
  • the decoder performs attribute decoding on the nodes in the current layer according to the second region adaptive layered combined transform mode, and then determines the attribute reconstruction value of the nodes in the current layer
  • the target decoding mode is the third region adaptive layered combined transform mode
  • the decoder performs attribute decoding on the nodes in the current layer according to the third region adaptive layered combined transform mode, and then determines the attribute reconstruction value of the nodes in the current layer.
  • the code stream is parsed to determine the first syntax identification information; then, when the first syntax identification information indicates that the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode, the code stream is parsed to determine the target decoding mode of the current layer; finally, the decoding end uses the parsed target decoding mode to reconstruct the attributes of the point cloud, thereby improving the decoding efficiency of the point cloud attributes, thereby improving Point cloud decoding performance.
  • the decoding method further includes:
  • the step of parsing the code stream and determining the target decoding mode of the current layer is not performed.
  • the index value of the current layer is determined in the following two ways:
  • Method 1 Determine the index value of the current layer according to the seventh syntax identification information
  • the implementation of method 1 may include:
  • the code stream is parsed to determine a value of seventh syntax identification information; the seventh syntax identification information is used to indicate an index value of a current layer.
  • the index value of the current layer is 1.
  • Method 2 The decoding end and the encoding end use the same process to determine the index value of the current layer. In this way, the encoding end does not need to transmit the codeword indicating the index value of the current layer to the decoding end, and the decoding end does not need to parse the corresponding codeword, which can also improve the decoding efficiency to a certain extent.
  • the number of layers included in the current sequence can be represented as attr_code_mode_cnt. It should be noted that the number of layers is the number of decoding layers corresponding to the inter-frame prediction mode and/or intra-frame prediction mode that can be adaptively selected in the current sequence.
  • Attr_code_mode_cnt is 10.
  • the index value of the current layer can be expressed as i, where i is an integer greater than or equal to 0.
  • the ninth value is 0.
  • the step of parsing the code stream and determining the target decoding mode of the current layer is performed, which can be expressed as:
  • the implementation of performing attribute decoding on the nodes in the current layer according to the target decoding mode and determining the attribute reconstruction value of the nodes in the current layer in S103 may include S1031 to S1034:
  • the decoder determines the attribute prediction value of the node in the current layer according to the number of adjacent nodes of the node in the current layer.
  • linear fitting can be performed using the reconstructed attributes of the neighboring nodes of the nodes in the current layer and the geometric distance of each neighboring node from the current node to obtain the predicted attribute values of the nodes in the current layer.
  • the attribute prediction for the nodes in the current layer can be based on intra-frame attribute prediction transformation or inter-frame attribute prediction transformation, which is not specifically limited here.
  • the implementation of determining the attribute prediction value of the node in the current layer in S1031 may include:
  • Linear fitting is performed based on the attribute reconstruction values corresponding to the adjacent nodes and the geometric distances between the nodes in the current layer and the adjacent nodes to determine the attribute prediction values of the nodes in the current layer.
  • the regional adaptive hierarchical transformation mode is a Haar wavelet transform, which can transform the point cloud attribute information from the spatial domain to the frequency domain, further reducing the correlation between the point cloud attributes.
  • the main idea is to transform the nodes in each layer from the three dimensions of x, y, and z in a bottom-up manner according to the octree structure, and iterate until the root node of the octree.
  • the basic idea is to perform wavelet transform based on the hierarchical structure of the octree, associate the attribute information with the octree nodes, and recursively transform the attributes of the occupied nodes in the same parent node in a bottom-up manner, and transform the nodes in each layer from the three dimensions of x, y, and z until they are transformed to the root node of the octree.
  • the nodes of the same layer obtained after the transformation are transformed.
  • the first coefficient is passed to the node of the next layer for further transformation, and all the second coefficients are decoded and determined by the arithmetic decoder.
  • the forward transform is a RAHT forward transform
  • the first coefficient value is a DC coefficient
  • the second coefficient prediction value is an AC coefficient prediction value
  • determining the second coefficient prediction value of the node in the current layer includes:
  • the first intermediate prediction value and the second intermediate prediction value of the node in the current layer are added to obtain the second coefficient prediction value of the node in the current layer.
  • the first intermediate prediction value can be expressed as w1*predIntraVal
  • the second intermediate prediction value can be expressed as w2*predIntraVal
  • forward transforming the nodes in the current layer according to the regional adaptive hierarchical combined transformation mode to determine the first intermediate prediction value and the second intermediate prediction value of the nodes in the current layer includes:
  • the second attribute prediction value of the node in the current layer is multiplied by the second target weight to obtain a second intermediate prediction value of the node in the current layer.
  • the RAHT intra-frame prediction value of the current node is predIntraVal (first attribute prediction value)
  • the inter-frame prediction value is predInterVal (second attribute prediction value)
  • the first target weight can be expressed as w1
  • the second target weight can be expressed as w2.
  • determining the first target weight and the second target weight of the current layer may include:
  • the target weight combination includes a first target weight and a second target weight.
  • the weight index value may be in parameter form or in digital form, and the embodiment of the present application does not impose any limitation on this.
  • the weight index value may be in digital form, such as a weight index value of 2.
  • the target weight combination includes a first target weight and a second target weight. In another embodiment, the target weight combination includes a first target weight, a second target weight and a third target weight.
  • the number of target weights included in the target weight combination is related to the target decoding mode.
  • the target weight combination includes a first target weight w1; if the target decoding mode is a first region adaptive layered inter-frame transform mode, the target weight combination includes a second target weight w2; if the target decoding mode is a second region adaptive layered inter-frame transform mode, the target weight combination includes a third target weight w3; if the target decoding mode is a first region adaptive layered combined transform mode, the target weight combination includes a first target weight w1 and a second target weight w2; if the target decoding mode is a second region adaptive layered combined transform mode, the target weight combination includes a first target weight w1 and a third target weight w3; if the target decoding mode is a third region adaptive layered combined transform mode, the target weight combination includes a first target weight w1, a second target weight w2 and a third target weight w3.
  • the target decoding mode of the front layer is the regional adaptive layered combined transform mode
  • the inter-frame prediction values and intra-frame prediction values of different RAHT transform layers will be merged, and the best prediction value will be finally obtained according to different weights, thereby further improving the RAHT decoding efficiency of point cloud attributes.
  • the method when the target decoding mode of the current layer is a region adaptive hierarchical inter-frame transform mode or a region adaptive hierarchical combined transform mode, after determining the second coefficient prediction value of the node in the current layer, the method further includes:
  • the node in the current layer is forward transformed using the regional adaptive hierarchical intra transform mode to obtain the intermediate second coefficient prediction value of the node in the current layer;
  • the intermediate second coefficient prediction value is used as the second coefficient prediction value of the node in the current layer.
  • the tenth value is 0.
  • any prediction decoding mode it is first determined whether the attribute prediction value between frames is equal to zero. If it is not equal to zero, the current prediction value will be directly used as the prediction value of the AC coefficient of the current node. Otherwise, the AC coefficient obtained by intra-frame prediction will be used as the AC coefficient prediction value of the current node.
  • the second coefficient value is also referred to as an AC coefficient reconstruction value.
  • determining the second coefficient value corresponding to the node of the current layer according to the second coefficient prediction value includes:
  • the second coefficient values of the nodes in the current layer are determined according to the second coefficient prediction values and the second coefficient inverse residual values corresponding to the nodes in the current layer.
  • the first coefficient may refer to a low-frequency coefficient, which may also be called a direct current (DC) coefficient;
  • the second coefficient may refer to a high-frequency coefficient, which may also be called an alternating current (AC) coefficient.
  • DC direct current
  • AC alternating current
  • g′ L,2x,y,z and g′ L,2x+1,y,z are two attribute DC coefficients of neighboring points in the L layer.
  • the information of the L-1 layer is the AC coefficient f′ L-1,x,y,z and the DC coefficient g′ L-1,x,y,z ; then, f′ L-1,x,y,z will no longer be transformed and will be directly quantized and decoded, and g′ L-1,x,y,z will continue to look for neighbors for transformation.
  • the weights (the number of non-empty child nodes in the node) corresponding to g′ L,2x,y,z and g′ L,2x+2,y ,z are w′ L,2x,y,z and w′ L,2x+1,y,z (abbreviated as w′ 0 and w′ 1 ) respectively, and the weight of g′ L-1,x,y,z is w′ L-1,x,y,z .
  • the general transformation formula is:
  • T w0,w1 is a transformation matrix, and the transformation matrix will be updated as the weights corresponding to each point change adaptively.
  • the forward transformation of RAHT (also referred to as "RAHT forward transformation") is shown in the aforementioned FIG. 35A.
  • the inverse RAHT transform is performed based on the DC coefficient and AC coefficient of the point in the current slice, so that the attribute reconstruction value of the point in the current slice can be restored.
  • the inverse RAHT transform (also referred to as "RAHT inverse transform” or "RAHT inverse transform”) is shown in FIG. 35B .
  • the implementation steps of the decoding end are as follows:
  • the attribute prediction value of each child node is used to perform RAHT transformation to obtain the corresponding DC coefficient and AC coefficient.
  • the AC coefficient of the predicted node and the AC coefficient parsed in the bitstream are used to restore the AC coefficient of the current node.
  • the AC coefficient and DC coefficient of the current node are used to perform an inverse RAHT transform, thereby recovering the attribute reconstruction value of each child node of the current node.
  • the decoding method further includes:
  • a preset threshold is used to determine whether a node in the current layer is allowed to perform attribute prediction.
  • the preset threshold is a value that is set.
  • the preset threshold can be a value agreed upon by both the decoder and the encoder.
  • the preset threshold can also be determined by the decoder by parsing the bitstream. The embodiment of the present application does not impose any restrictions on the method for obtaining the preset threshold.
  • the attribute information of the neighboring nodes of the current layer is continued to be determined. otherwise, if the number of adjacent nodes in the current layer is less than the preset threshold, it can be determined that the current layer does not allow attribute prediction. In this case, the attribute prediction of the nodes in the current layer is directly stopped, and the attribute prediction of the next layer can be performed.
  • the decoding method further includes:
  • the neighboring nodes of a node include at least: neighboring nodes coplanar with the node and neighboring nodes colinear with the node.
  • the spatial position information of each node in the current layer may be the position information of the node, specifically the three-dimensional coordinate information (x, y, z).
  • the neighboring nodes of a node may include: neighboring nodes coplanar with the node and neighboring nodes colinear with the node.
  • a grid filling block may represent the current node
  • a slash filling block may represent some neighboring nodes coplanar and colinear with the current node.
  • the decoding method further includes:
  • the parent node neighbor nodes of each node are determined; wherein the parent node neighbor nodes of the node at least include: neighbor nodes coplanar with the parent node of the node and neighbor nodes colinear with the parent node of the node.
  • determining the number of adjacent nodes of the current layer includes:
  • RAHT can be used as both a transform and a prediction, resulting in high complexity.
  • the relevant technology sets a start condition for whether the current node is allowed to perform attribute prediction, specifically: judging whether the number of adjacent nodes in the current layer is greater than a preset threshold. In this way, by setting the judgment condition for whether the current layer starts attribute prediction, the memory occupancy of point cloud attribute decoding can be reduced while ensuring complexity, and the decoding efficiency of the point cloud can also be improved.
  • the first grammar identification information is the first value
  • the fourth grammar identification information is used to indicate whether the node of the current layer is allowed to perform inter-frame prediction
  • the fifth grammar identification information is used to indicate whether the node of the current layer is allowed to perform intra-frame prediction
  • the first grammar identification information is the second value.
  • the first syntax element information is the first value only when the fourth syntax identification information and the fifth syntax identification information are both the first value, that is, when it is determined that the node for the current attribute decoding allows inter-frame prediction and intra-frame prediction, the first syntax element information is the first value.
  • the decoding method further includes:
  • the first syntax identification information indicates that the current layer does not allow adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode, parse the bitstream to determine the fourth syntax identification information;
  • the attributes of the nodes in the current layer are decoded according to the region adaptive hierarchical intra-frame transform mode to determine the attribute reconstruction values of the nodes in the current layer.
  • the fourth syntax identification information may be represented as !disableAttrInterPred
  • the fifth syntax identification information may be represented as raht_prediction_enabled.
  • !disableAttrInterPred it is determined that the node of the current layer is allowed to perform inter-frame prediction; when !disableAttrInterPred is false, it is determined that the node of the current layer is not allowed to perform inter-frame prediction.
  • raht_prediction_enabled when raht_prediction_enabled is true (1 or true), it is determined that the nodes of the current layer are allowed to perform intra-frame prediction; when raht_prediction_enabled is false (0 or false), it is determined that the nodes of the current layer are not allowed to perform intra-frame prediction.
  • the fourth grammar identification information and the fifth grammar identification information may be high-level grammar elements, and the fourth grammar identification information and the fifth grammar identification information may be set in an attribute parameter set (aps).
  • the decoder determines the attribute parameter set by parsing the bitstream; and determines the fourth syntax identification information and the fifth syntax identification information corresponding to the current layer from the attribute parameter set.
  • the decoder determines the target decoding mode corresponding to the current layer by parsing the bitstream. In other words, the decoder will continue to parse the bitstream to determine the target attribute of the current layer only when it is determined that the first syntax identification information corresponding to the current layer is true (1 or true).
  • the decoding method further includes: when the fourth syntax identification information indicates that the nodes in the current layer allow inter-frame prediction, performing attribute decoding on the nodes in the current layer according to the region adaptive hierarchical inter-frame transform mode, determining the nodes in the current layer. The attribute reconstruction value of the point.
  • the decoding method further includes:
  • the value of the fourth syntax identification information is the first value, it is determined that the node of the current layer allows inter-frame prediction
  • the value of the fourth syntax identification information is the second value, it is determined that the nodes of the current layer are not allowed to perform inter-frame prediction.
  • the decoding method further includes:
  • the value of the fifth syntax identification information is the first value, it is determined that the node of the current layer allows inter-frame prediction
  • the value of the fifth syntax identification information is the second value (false)
  • the decoding method further includes:
  • the sixth syntax identification information can be expressed as attr_coding_type.
  • a decoding method is provided, which is applied to a decoder.
  • the decoder parses the bitstream and determines the first syntax identification information; then, when the first syntax identification information indicates that the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode, the decoder parses the bitstream and determines the target decoding mode of the current layer; finally, the decoder uses the parsed target decoding mode to reconstruct the attributes of the point cloud, thereby improving the decoding efficiency of the point cloud attributes, and further improving the decoding performance of the point cloud.
  • FIG. 42 a schematic diagram of a flow chart of an encoding method provided in an embodiment of the present application is shown. As shown in FIG. 42 , the method may include S301 to S302:
  • the encoding method is applied to a point cloud encoder (hereinafter referred to as "encoder").
  • the encoding method may be a point cloud attribute encoding method, and more specifically, may be a point cloud attribute RAHT transform prediction adaptively selecting inter-frame prediction or intra-frame prediction for encoding.
  • the main thing here is to introduce the corresponding target coding mode for each layer in the current sequence in the attribute block header information parameter set (ABH), and the corresponding target coding mode can be adaptively selected for each layer (Layer), thereby improving the coding efficiency of point cloud attributes.
  • the current layer may be one of the layers in the current video frame.
  • the current layer includes at least one node.
  • the current layer may be referred to as a current attribute coding layer, a current coding layer, a current slice, etc.
  • the embodiment of the present application does not impose any limitation on this.
  • the current layer is a coding layer obtained by downsampling along the first direction, the second direction and the third direction, wherein the first direction is the z-axis direction, the second direction is the y-axis direction, and the third direction is the x-axis direction.
  • the embodiment of the present application does not limit the order of the first direction, the second direction, and the third direction.
  • it can be the second direction, the first direction, and the third direction, or it can be the third direction, the second direction, and the first direction.
  • the current layer is not limited to a coding layer obtained by downsampling once along the first direction, the second direction, and the third direction.
  • the current layer may also be multiple coding layers obtained by downsampling once along the first direction, the second direction, and the third direction.
  • the current layer may also be a layer composed of at least one node in a coding layer. The embodiment of the present application does not impose any limitation on this.
  • the first syntax identification information is used to indicate that the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode.
  • the encoding method further includes: determining first grammar identification information.
  • the implementation of determining the first grammar identification information may include:
  • the value of the first syntax identification information is determined to be a first value; the current coefficient group includes at least one layer, and the current layer is one of the at least one layer;
  • the value of the first syntax identification information is determined to be the second value.
  • the implementation of determining the first grammar identification information may further include:
  • the value of the first syntax identification information is determined to be the second value.
  • the first value is different from the second value, and the first value and the second value can be in parameter form or in digital form.
  • the first syntax identification information can be a parameter written in the profile or a flag value, which is not specifically limited here.
  • the first value can be set to 1 and the second value can be set to 0; or, the first value can be set to 0 and the second value can be set to 1; or, the first value can be set to true and the second value can be set to false; or, the first value can be set to false and the second value can be set to true; but this is not specifically limited here.
  • the first grammar identification information acts as a switch, that is, when the first grammar identification information is a first value (such as 1 or true), it indicates that the encoding algorithm of the embodiment of the present application is started, that is, the encoding algorithm of the embodiment of the present application is executed; when the first grammar identification information is a second value (such as 0 or false), it indicates that the encoding algorithm of the embodiment of the present application is not started, that is, the encoding algorithm of the embodiment of the present application is not executed.
  • a first value such as 1 or true
  • different layers can adaptively select intra-frame prediction and/or inter-frame prediction to perform attribute coding on the nodes, which can fully consider the distribution of AC coefficients of different layers, thereby improving the coding efficiency of RAHT.
  • the value of the sixth grammar identification information is determined; the sixth grammar identification information is used to indicate whether the nodes of the current layer are allowed to perform attribute prediction; the sixth grammar identification information is encoded, and the obtained encoded bits are written into the bit stream.
  • determining the value of the sixth grammar identification information may include: if it is determined that the nodes of the current layer allow attribute prediction, setting the value of the sixth grammar identification information to a first value; if it is determined that the nodes of the current layer do not allow attribute prediction, setting the value of the sixth grammar identification information to a second value.
  • the decoder can subsequently parse the bitstream and make a judgment based on the value of the sixth syntax identification information.
  • the encoder and the decoder can determine whether the node of the current layer is allowed to perform attribute prediction by a judgment method agreed upon by both parties.
  • the implementation of determining whether the node of the current layer is allowed to perform attribute prediction may include:
  • the encoder and decoder use an agreed judgment method to determine whether the nodes in the current layer are allowed to perform attribute prediction. In this way, the encoder does not need to write the codewords encoded by the sixth grammatical identification information into the bitstream, which can save codewords and thus improve coding efficiency.
  • the target coding mode can be expressed as attr_code_mode[i]; where i is the index value of the current layer.
  • index value i is assigned only when the current layer satisfies the three conditions of allowing attribute prediction, allowing inter-frame prediction, and allowing intra-frame prediction.
  • the encoder directly skips the current layer, directly performs attribute encoding of the next layer, and assigns index 2 to the next layer; if the current layer meets the three conditions of allowing attribute prediction, allowing inter-frame prediction, and allowing intra-frame prediction, the encoder performs attribute encoding on the current layer, adds 1 to the index value (i++), obtains the updated index value (3), and passes the index value 3 to the next layer.
  • the target coding mode includes a region adaptive hierarchical intra-frame transform mode, a region adaptive hierarchical inter-frame transform mode and a region adaptive hierarchical combined transform mode;
  • the regional adaptive hierarchical intra-frame transform mode representation adopts the intra-frame prediction mode to perform attribute prediction transform coding on the nodes of the current layer; the regional adaptive hierarchical inter-frame transform mode representation adopts the inter-frame prediction mode to perform attribute prediction transform coding on the nodes of the current layer; the regional adaptive hierarchical combined transform mode representation adopts the intra-frame prediction mode combined with the inter-frame prediction mode to perform attribute prediction transform coding on the nodes of the current layer.
  • the regional adaptive hierarchical inter-frame transform mode includes a first regional adaptive hierarchical inter-frame transform mode and a second regional adaptive hierarchical inter-frame transform mode;
  • the regional adaptive hierarchical combined transform mode includes a first regional adaptive hierarchical combined transform mode, a second regional adaptive hierarchical combined transform mode and a third regional adaptive hierarchical combined transform mode;
  • the first region adaptive hierarchical inter-frame transform mode represents the method of using the geometric information of the node to determine the co-located prediction node to perform attribute prediction transform coding on the nodes of the current layer;
  • the second region adaptive hierarchical combination transform mode characterization uses the reference frame cache to determine the co-location prediction node for the current layer. Points are subjected to attribute prediction transform coding;
  • the first region adaptive hierarchical combined transform mode characterization uses a combined region adaptive hierarchical intra-frame transform mode and a first region adaptive hierarchical inter-frame transform mode to perform attribute prediction transform coding on nodes of the current layer;
  • the second region adaptive hierarchical combined transform mode characterization uses a combined region adaptive hierarchical intra-frame transform mode and a second region adaptive hierarchical inter-frame transform mode to perform attribute prediction transform coding on the nodes of the current layer;
  • the third region adaptive hierarchical combined transform mode characterization uses a combined region adaptive hierarchical intra-frame transform mode, a first region adaptive hierarchical inter-frame transform mode, and a second region adaptive hierarchical inter-frame transform mode to perform attribute prediction transform coding on nodes of the current layer.
  • the region adaptive layered intra-frame transform mode the first region adaptive layered inter-frame transform mode, the second region adaptive layered inter-frame transform mode, the first region adaptive layered combined transform mode, the second region adaptive layered combined transform mode and the third region adaptive layered combined transform mode, please refer to the relevant description on the decoding end in the previous text and will not be repeated here.
  • the encoding method further includes:
  • the second syntax identification information is added to the attribute block header information parameter set, and the attribute block header information parameter set is coded, and the obtained coded bits are written into the bitstream.
  • the encoder determines the second syntax identification information according to the value of the target coding mode, wherein the second syntax identification information is used to indicate the target coding mode of the current layer.
  • the value of the second syntax identification information may be in parameter form or in digital form, and the embodiment of the present application does not impose any limitation on this.
  • the encoding method further includes:
  • the target coding mode of the nodes in the current layer is encoded, and the obtained coded bits are written into the bitstream.
  • the encoding end may adopt a rate-distortion algorithm to obtain a cost value corresponding to at least one candidate coding mode, and then determine the target coding mode based on the cost value corresponding to at least one candidate coding mode, and determine the first grammar identification information; the first grammar identification information is used to indicate whether the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode, and then the first grammar identification information and the second grammar identification information are written into the bitstream.
  • the encoding end may also directly write the second grammar identification information into the bitstream, so that on the subsequent decoding side, the decoding end can directly decode the second grammar identification information to obtain the target decoding mode. This application does not impose any limitation on this.
  • the method may also include: adding the target coding mode of the current layer to the attribute block header information parameter set; encoding the attribute block header information parameter set, and writing the obtained coded bits into the bitstream.
  • the method may also include: dividing the current sequence into coding layers, determining at least one coding layer; and adding all the target coding modes corresponding to the at least one coding layer to the attribute block header information parameter set.
  • the decoding end can directly decode the target decoding mode of the current slice from the attribute block header information parameter set.
  • determining the implementation of the second grammar identification information according to the target coding mode may include:
  • the value of the second syntax identification information is determined according to a target decoding mode of the current layer adopting an inter-frame prediction mode and/or an intra-frame prediction mode.
  • the implementation of determining the value of the second syntax identification information according to the target decoding mode of the current layer using the inter-frame prediction mode and/or the intra-frame prediction mode may include:
  • the target coding mode of the current layer is the region adaptive hierarchical intra transform mode, determining that the value of the second syntax identification information is a third value
  • the target coding mode of the current layer is the first region adaptive hierarchical inter-frame transform mode, determining the value of the second syntax identification information to be a fourth value;
  • the target coding mode of the current layer is the second region adaptive hierarchical inter-frame transform mode, determining the value of the second syntax identification information to be a fifth value;
  • the target coding mode of the current layer is the first region adaptive hierarchical combined transform mode, determining the value of the second syntax identification information to be a sixth value;
  • the target coding mode of the current layer is the second region adaptive hierarchical combined transform mode, determining the value of the second syntax identification information to be the seventh value;
  • the value of the second syntax identification information is determined to be the eighth value.
  • the third value, the fourth value, the fifth value, the sixth value, the seventh value and the eighth value are different.
  • the third value, the fourth value, the fifth value, the sixth value, the seventh value, and the eighth value may be in parameter form or in digital form.
  • the third value is 0, the fourth value is 1, the fifth value is 2, the sixth value is 3, the seventh value is 4, and the eighth value is 5.
  • the present application embodiment does not impose any restrictions on the settings of the third value, the fourth value, the fifth value, the sixth value, the seventh value, and the eighth value.
  • S302 Perform attribute encoding on the nodes in the current layer according to the target coding mode to determine the attribute reconstruction values of the nodes in the current layer.
  • the encoder can perform attribute encoding on the nodes in the current layer according to the target coding mode, and then determine the attribute reconstruction values of the nodes in the current layer.
  • the encoder performs attribute encoding on the nodes in the current layer according to the region adaptive layered intra-frame transform mode, and then determines the attribute reconstruction values of the nodes in the current layer; if the target coding mode is a first region adaptive layered inter-frame transform mode, the encoder performs attribute encoding on the nodes in the current layer according to the first region adaptive layered inter-frame transform mode, and then determines the attribute reconstruction values of the nodes in the current layer; if the target coding mode is a second region adaptive layered inter-frame transform mode, the encoder performs attribute encoding on the nodes in the current layer according to the second region adaptive layered inter-frame transform mode, and then determines the attribute reconstruction values of the nodes in the current layer.
  • the encoder performs attribute encoding on the nodes in the current layer according to the first region adaptive layered combined transform mode, and then determines the attribute reconstruction value of the nodes in the current layer; if the target coding mode is the second region adaptive layered combined transform mode, the encoder performs attribute encoding on the nodes in the current layer according to the second region adaptive layered combined transform mode, and then determines the attribute reconstruction value of the nodes in the current layer; if the target coding mode is the third region adaptive layered combined transform mode, the encoder performs attribute encoding on the nodes in the current layer according to the third region adaptive layered combined transform mode, and then determines the attribute reconstruction value of the nodes in the current layer.
  • the encoder determines the first grammatical identification information when it is determined that the nodes of the current layer allow attribute prediction; when the first grammatical identification information indicates that the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode, the encoder determines the target coding mode of the current layer; attributes of the nodes in the current layer are encoded according to the target coding mode, and the attribute reconstruction values of the nodes in the current layer are determined, thereby improving the coding efficiency of the point cloud attributes, and then improving the coding performance of the point cloud.
  • the encoding method further includes:
  • the third syntax identification information is used to indicate the number of layers included in the current sequence where the current layer is located;
  • the step of determining the target coding mode of the current layer is not performed
  • the third syntax identification information is encoded, and the obtained encoded bits are written into the bitstream.
  • the number of layers included in the current sequence can be represented as attr_code_mode_cnt. It should be noted that the number of layers is the number of coding layers corresponding to the inter-frame prediction mode and/or intra-frame prediction mode that can be adaptively selected in the current sequence.
  • Attr_code_mode_cnt is 10.
  • the index value of the current layer can be expressed as i, where i is an integer greater than or equal to 0.
  • the ninth value is 0.
  • the step of parsing the bitstream and determining the target coding mode of the current layer is performed, which can be expressed as:
  • the implementation of determining the target coding mode of the current layer in S301 may include S3011 to S3013:
  • At least one candidate encoding mode includes at least one of the following: a region adaptive layered intra-frame transform mode, a first region adaptive layered inter-frame transform mode, a second region adaptive layered inter-frame transform mode, a first region adaptive layered combined transform mode, a second region adaptive layered combined transform mode, and a third region adaptive layered combined transform mode.
  • a cost calculation is performed for at least one candidate coding mode.
  • the cost here may refer to a distortion value, a rate-distortion cost value, or other cost values, which are not specifically limited here.
  • implementation of S3012 may include: performing rate-distortion cost calculation according to a precoding result of each of at least one candidate coding mode, and determining a cost value of each of the at least one candidate coding mode.
  • cost calculation is performed based on the precoding result of each of at least one candidate coding mode to determine the cost value of each of at least one candidate coding mode, which may include: performing rate-distortion cost calculation based on the precoding result of each of at least one candidate coding mode to determine the cost value of each of at least one candidate coding mode.
  • J represents the rate distortion cost value
  • R represents the bitstream required to be encoded by the candidate coding mode
  • can be calculated by the attribute quantization parameter.
  • the current calculation method of ⁇ is as follows:
  • QP represents a quantization parameter
  • N can be set to different values according to reflectivity and color.
  • the implementation of S3023 may include:
  • the candidate coding mode corresponding to the minimum cost value is determined as the target coding mode of the current layer.
  • At least one candidate coding mode includes: a second region adaptive layered combined transform mode and a third region adaptive layered combined transform mode
  • cost calculation is performed for the second region adaptive layered combined transform mode and the third region adaptive layered combined transform mode, respectively, to determine the first generation value of the second region adaptive layered combined transform mode and the second generation value of the third region adaptive layered combined transform mode; based on the first generation value and the second generation value, determine the target coding mode of the current layer.
  • the target coding mode of the current layer is determined based on the first generation value and the second generation value. Specifically, it can be: if the first generation value is less than the second generation value, the second region adaptive layered combined transform mode is determined as the target coding mode of the current layer; or, if the first generation value is greater than the second generation value, the third region adaptive layered combined transform mode is determined as the target coding mode of the current layer.
  • the second region adaptive layered combined transform mode can be determined as the target coding mode of the current layer, or the third region adaptive layered combined transform mode can be determined as the target coding mode of the current layer, without specific limitation here.
  • the implementation of performing attribute encoding on the nodes in the current layer according to the target coding mode and determining the attribute reconstruction value of the nodes in the current layer in S302 may include S3021 to S3024:
  • the encoder determines the attribute prediction value of the node in the current layer according to the number of adjacent nodes of the node in the current layer.
  • linear fitting can be performed using the reconstructed attributes of the neighboring nodes of the nodes in the current layer and the geometric distance of each neighboring node from the current node to obtain the predicted attribute values of the nodes in the current layer.
  • the attribute prediction for the nodes in the current layer can be based on intra-frame attribute prediction transformation or inter-frame attribute prediction transformation, which is not specifically limited here.
  • the implementation of determining the attribute prediction value of the node in the current layer may include:
  • Linear fitting is performed based on the attribute reconstruction values corresponding to the adjacent nodes and the geometric distances between the nodes in the current layer and the adjacent nodes to determine the attribute prediction values of the nodes in the current layer.
  • the 19 neighboring nodes of the node in the current layer are first determined, and then the spatial geometric distance between the neighboring nodes and each node of the current node is used to perform linear weighted prediction on the attributes of each node, and finally the attribute prediction value of each node is determined based on the linear weighted prediction value.
  • the regional adaptive hierarchical transformation mode is a Haar wavelet transform, which can transform the point cloud attribute information from the spatial domain to the frequency domain, further reducing the correlation between the point cloud attributes.
  • the main idea is to transform the nodes in each layer from the three dimensions of x, y, and z in a bottom-up manner according to the octree structure, and iterate until the root node of the octree.
  • the basic idea is to perform wavelet transform based on the hierarchical structure of the octree, associate the attribute information with the octree nodes, and recursively transform the attributes of the occupied nodes in the same parent node in a bottom-up manner, and transform the nodes in each layer from the three dimensions of x, y, and z until they are transformed to the root node of the octree.
  • the nodes of the same layer obtained after the transformation are transformed.
  • the first coefficient is passed to the node of the next layer for further transformation, and all the second coefficients are encoded by the arithmetic encoder.
  • the forward transform is a RAHT forward transform
  • the first coefficient value is a DC coefficient
  • the second coefficient prediction value is an AC coefficient prediction value
  • determining the second coefficient prediction value of the node in the current layer includes:
  • the first intermediate prediction value and the second intermediate prediction value of the node in the current layer are added to obtain the second coefficient prediction value of the node in the current layer.
  • the first intermediate prediction value can be expressed as w1*predIntraVal
  • the second intermediate prediction value can be expressed as w2*predIntraVal
  • forward transforming the nodes in the current layer according to the regional adaptive hierarchical combined transformation mode to determine the first intermediate prediction value and the second intermediate prediction value of the nodes in the current layer may include:
  • the target weight combination includes a first target weight and a second target weight
  • the second attribute prediction value of the node in the current layer is multiplied by the second target weight to obtain a second intermediate prediction value of the node in the current layer.
  • the weight index value may be in parameter form or in digital form, and the embodiment of the present application does not impose any limitation on this.
  • the weight index value may be in digital form, such as a weight index value of 2.
  • the target weight combination includes a first target weight and a second target weight. In another embodiment, the target weight combination includes a first target weight, a second target weight and a third target weight.
  • the number of target weights included in the target weight combination is related to the target coding mode.
  • the target weight combination includes a first target weight w1; if the target coding mode is a first region adaptive layered inter-frame transform mode, the target weight combination includes a second target weight w2; if the target coding mode is a second region adaptive layered inter-frame transform mode, the target weight combination includes a third target weight w3; if the target coding mode is a first region adaptive layered combined transform mode, the target weight combination includes a first target weight w1 and a second target weight w2; if the target coding mode is a second region adaptive layered combined transform mode, the target weight combination includes a first target weight w1 and a third target weight w3; if the target coding mode is a third region adaptive layered combined transform mode, the target weight combination includes a first target weight w1, a second target weight w2 and a third target weight w3.
  • the target coding mode of the current layer is the regional adaptive layered combined transform mode
  • the inter-frame prediction values and intra-frame prediction values of different RAHT transform layers will be merged, and the best prediction value will be finally obtained according to different weights, thereby further improving the RAHT coding efficiency of point cloud attributes.
  • the target weight combination is encoded with the corresponding weight index value in the preset weight table, and the obtained encoded bits are written into the bitstream.
  • the method when the target coding mode of the current layer is a region adaptive hierarchical inter-frame transform mode or a region adaptive hierarchical combined transform mode, after determining the second coefficient prediction value of the node in the current layer, the method further includes:
  • the second coefficient prediction value of the node in the current layer is the tenth value, forward transforming the node in the current layer according to the region adaptive hierarchical intra transform mode to obtain an intermediate second coefficient prediction value of the node in the current layer;
  • the intermediate second coefficient prediction value is used as the second coefficient prediction value of the node in the current layer.
  • the first coefficient may refer to a low-frequency coefficient, also known as a direct current (DC) coefficient;
  • the second coefficient may refer to a high-frequency coefficient, also known as an alternating current (AC) coefficient.
  • DC direct current
  • AC alternating current
  • the implementation of determining the second coefficient value corresponding to the node of the current layer according to the second coefficient prediction value may include:
  • the second coefficient values of the nodes in the current layer are determined according to the second coefficient prediction values and the second coefficient inverse residual values corresponding to the nodes in the current layer.
  • the implementation of determining the second coefficient encoding residual value of the node in the current layer may include:
  • the second coefficient residual value is quantized to obtain the second coefficient quantized residual value of the node in the current layer.
  • the encoding method further includes:
  • the second coefficient quantization residual value of the node in the current layer is encoded, and the obtained encoding bits are written into the bit stream.
  • the first coefficient may refer to a low-frequency coefficient, also known as a direct current (DC) coefficient;
  • the second coefficient may refer to a high-frequency coefficient, also known as an alternating current (AC) coefficient.
  • DC direct current
  • AC alternating current
  • g′ L,2x,y,z and g′ L,2x+1,y,z are two attribute DC coefficients of neighboring points in the L layer.
  • the information of the L-1 layer is the AC coefficient f′ L-1,x,y,z and the DC coefficient g′ L-1,x,y,x ; then, f′ L-1,x,y,z will no longer be transformed and will be directly quantized and encoded, and g′ L-1,x,y,z will continue to look for neighbors for transformation.
  • the weights (the number of non-empty child nodes in the node) corresponding to g′ L,2x,y,z and g′ L,2x+2,y ,z are w′ L,2x,y,z and w′ L,2x+1,y,z (abbreviated as w′ 0 and w′ 1 ) respectively, and the weight of g′ L-1,x,y,z is w′ L-1,x,y,z .
  • the general transformation formula is:
  • T w0,w1 is a transformation matrix, and the transformation matrix will be updated as the weights corresponding to each point change adaptively.
  • the forward transformation of RAHT (also referred to as "RAHT forward transformation") is shown in the aforementioned FIG. 35A.
  • the inverse RAHT transform is performed according to the DC coefficient and AC coefficient of the point in the current slice, so that the attribute reconstruction value of the point in the current slice can be restored.
  • the inverse RAHT transform (also referred to as "RAHT inverse transform” or "RAHT inverse transform”) is shown in the aforementioned FIG. 35B.
  • the implementation steps of the encoding end are as follows:
  • the attribute prediction value of each child node is used to perform RAHT transformation to obtain the corresponding DC coefficient and AC coefficient.
  • the attributes of each child node of the current node are transformed through RAHT transformation to obtain DC coefficient and AC coefficient;
  • the predicted value of the AC coefficient obtained by the prediction node is used to predict the AC of the current node, and finally the AC prediction residual coefficient of each child node is quantized and encoded.
  • the AC reconstruction coefficient of the current node is restored using the dequantized value of the AC prediction residual coefficient and the predicted value of the AC coefficient.
  • the AC coefficient and DC coefficient of the current node are used to perform an inverse RAHT transform to restore the attribute reconstruction value of each child node of the current node.
  • the encoder determines the first grammatical identification information; when the first grammatical identification information indicates that the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode, the encoder determines the target coding mode of the current layer; attributes of the nodes in the current layer are encoded according to the target coding mode, and attribute reconstruction values of the nodes in the current layer are determined, thereby improving the coding efficiency of the point cloud attributes, and further improving the coding performance of the point cloud.
  • the encoding method further includes:
  • the fifth grammar identification information is used to indicate whether a node in the current layer is allowed to perform intra-frame prediction
  • the first syntax identification information is encoded, and the obtained encoded bits are written into a bit stream.
  • the first grammar element information is the first value, that is, when it is determined that the node of the current attribute encoding allows inter-frame prediction and intra-frame prediction, the first grammar element information is the first value.
  • the fourth syntax identification information may be represented as !disableAttrInterPred
  • the fifth syntax identification information may be represented as raht_prediction_enabled.
  • !disableAttrInterPred it is determined that the node of the current layer is allowed to perform inter-frame prediction; when !disableAttrInterPred is false, it is determined that the node of the current layer is not allowed to perform inter-frame prediction.
  • raht_prediction_enabled when raht_prediction_enabled is true (1 or true), it is determined that the nodes of the current layer are allowed to perform intra-frame prediction; when raht_prediction_enabled is false (0 or false), it is determined that the nodes of the current layer are not allowed to perform intra-frame prediction.
  • the encoding method further includes:
  • the value of the fourth grammar identification information is set to the first value
  • the value of the fourth syntax identification information is set to the second value
  • the fourth syntax identification information is encoded, and the obtained encoded bits are written into the bitstream.
  • the encoding method further includes:
  • the value of the fifth grammar identification information is set to the first value
  • the value of the fifth grammar identification information is set to the second value
  • the fifth syntax identification information is encoded, and the obtained encoded bits are written into the bitstream.
  • the encoding method further includes:
  • the sixth syntax identification information is used to indicate that the nodes in the current layer adopt the region adaptive layered inter-frame transform mode
  • the sixth syntax identification information is encoded, and the obtained encoded bits are written into the bitstream.
  • the encoding method further includes:
  • the encoding method further includes:
  • the neighboring nodes of a node include at least: neighboring nodes coplanar with the node and neighboring nodes colinear with the node.
  • the encoding method further includes:
  • the parent node neighbor nodes of each node are determined; wherein the parent node neighbor nodes of the node at least include: neighbor nodes coplanar with the parent node of the node and neighbor nodes colinear with the parent node of the node.
  • the spatial position information of each node in the current layer may be the position information of the node, specifically the three-dimensional coordinate information (x, y, z).
  • the neighboring nodes of a node may include: neighboring nodes coplanar with the node and neighboring nodes colinear with the node.
  • a grid filling block may represent the current node
  • a slash filling block may represent some neighboring nodes coplanar and colinear with the current node.
  • encoding and determining the number of adjacent nodes of the current layer includes:
  • RAHT can be used as both a transformation and a prediction, resulting in high complexity.
  • the related technology sets a start condition for whether the current node is allowed to perform attribute prediction, specifically: judging whether the number of adjacent nodes in the current layer is greater than a preset threshold. In this way, by setting the judgment condition for whether the current layer starts attribute prediction, the memory occupancy of point cloud attribute coding can be reduced while ensuring complexity, and the coding efficiency of the point cloud can also be improved.
  • the encoder determines the first syntax when determining that the node of the current layer allows attribute prediction. Identification information; when the first grammatical identification information indicates that the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode, the encoder determines the target coding mode of the current layer; attributes of the nodes in the current layer are encoded according to the target coding mode, and attribute reconstruction values of the nodes in the current layer are determined, thereby improving the coding efficiency of point cloud attributes, and further improving the coding performance of point cloud.
  • a point cloud attribute RAHT transform prediction RDO adaptively selects inter-frame prediction or intra-frame prediction encoding scheme.
  • the best encoding mode of each attribute encoding layer is adaptively selected according to the RDO method, and then the final best encoding mode is passed to the decoding end.
  • the decoding end uses the obtained best encoding mode to reconstruct the attributes of the point cloud, thereby further improving the encoding efficiency of the point cloud attributes.
  • a new coding scheme is introduced to improve the coding efficiency of point cloud attributes.
  • three attribute prediction coding schemes are combined: two inter-frame prediction coding schemes and one intra-frame prediction coding scheme.
  • the best coding mode of the current RAHT coding layer is obtained by using a rate-distortion optimization algorithm at the encoding end, namely: inter-frame prediction coding scheme 2 + intra-frame prediction coding, inter-frame prediction coding scheme 1 + inter-frame prediction coding scheme 2 + intra-frame prediction coding.
  • the best coding mode of the current RAHT coding layer is passed to the decoding end.
  • the decoding end uses the coding mode of the current layer RAHT to adaptively restore the AC coefficient of the current layer, thereby completing the entire attribute RAHT coding, and finally improving the RAHT attribute coding efficiency.
  • the RAHT attribute coding layer (i.e., the current layer) is first defined.
  • the current attribute RAHT transform coding order is to divide from the root node in sequence until it is divided to the voxel level (1x1x1), thereby completing the encoding and attribute reconstruction of the entire point cloud attribute.
  • the layer obtained by downsampling once along the Z direction, Y direction, and X direction is defined as a RAHT transform layer, i.e., layer.
  • a rate-distortion optimization algorithm is introduced to adaptively select the prediction coding method of the current layer, and two prediction coding modes are introduced: 1.
  • Intra-frame combined with inter-frame prediction coding mode 2 i.e., the second region adaptive hierarchical combined transform mode
  • Intra-frame prediction mode combined with inter-frame prediction coding mode 2 and inter-frame prediction coding mode 1 i.e., the third region adaptive hierarchical combined transform mode
  • the rate-distortion optimization algorithm is used to predict and encode the attribute information of the current layer node using two prediction modes at the encoding end, and finally the rate-distortion optimization algorithm is used to obtain the best encoding mode of the current layer, and the best encoding mode is passed to the decoding end, and the decoding end uses the predicted decoding mode obtained by analysis to reconstruct and restore the attribute information of the current layer point to be decoded.
  • the rate-distortion optimization algorithm the distortion D between the reconstructed attribute and the original attribute of each prediction mode is first calculated, and then the code stream R required for encoding of each prediction mode is obtained, and the rate-distortion cost calculation is shown in the above equations (36) and (37).
  • the target coding mode of each layer is finally added to the ABH parameter set.
  • the specific algorithm of the encoding end is as follows:
  • Step 1 Adaptively determine whether the nodes in the current layer can use attribute prediction based on the number of neighboring nodes in the current layer and the number of neighboring nodes of the parent node;
  • Step 2 If the nodes of the current layer can use attribute prediction and attribute inter-frame prediction, the rate-distortion optimization algorithm is introduced for the current layer. By encoding each node of the current layer, the cost corresponding to each prediction coding mode is calculated to obtain the optimal prediction coding mode.
  • Step 3 Finally, the best prediction coding mode is used to predict the attributes of the current layer nodes.
  • the specific algorithm of the decoding end is as follows:
  • Step 1 Adaptively determine whether the nodes in the current layer can use attribute prediction based on the number of neighboring nodes in the current layer and the number of neighboring nodes of the parent node;
  • Step 2 If the node of the current layer can adopt attribute prediction and can perform attribute inter-frame prediction, the node obtains the best prediction decoding mode of the current layer.
  • Step 3 Finally, the best prediction decoding mode is used to predict and decode the attributes of the current layer nodes.
  • the embodiment of the present application introduces a prediction coding mode (attr_code_mode[i]) in each RAHT coding layer when performing RAHT prediction coding on the attributes to adaptively select inter-frame prediction coding mode 2 combined with intra-frame prediction coding mode or intra-frame prediction coding combined with two inter-frame prediction coding modes, and finally passes the coding mode to the decoding end, which uses the coding mode to reconstruct the attributes of the point cloud.
  • the focus is on introducing a coding mode in each RAHT coding layer, obtaining the best coding mode by utilizing the rate-distortion optimization selection algorithm at the encoding end, and then using the decoding mode at the decoding end to reconstruct the attributes of the point cloud.
  • the coding mode of each layer is stored in ABH, and the decoding mode of the RAHT coding layer is obtained by ABH at the decoding end. There is no restriction on how the parameter is encoded.
  • the attribute inter-frame prediction mode can be further modified.
  • a coding mode is introduced to the current RAHT coding layer at the encoding end to represent which prediction coding mode is used to restore the AC coefficient of the current RAHT coding layer.
  • This scheme can further change the prediction coding mode to: inter-frame prediction coding mode 1, intra-frame prediction coding mode, inter-frame prediction coding mode 1, and determine the best coding mode for the current layer in the same way as the main scheme.
  • the decoding end also recovers the AC coefficient of the current layer based on the prediction coding mode of the current layer, thereby completing the entire RAHT attribute coding.
  • the attribute prediction mode can be further modified. Specifically, in the main scheme, for any prediction coding mode, first determine whether the attribute prediction value between frames is equal to zero. If it is not equal to zero, the current prediction value will be directly used as the prediction value of the AC coefficient of the current node, otherwise the AC coefficient obtained by intra-frame prediction will be used as the AC coefficient prediction value of the current node. In this scheme, the inter-frame prediction values and intra-frame prediction values of different RAHT transformation layers will be merged, and the best prediction value will be finally obtained according to different weights, so as to further improve the RAHT coding efficiency of point cloud attributes.
  • the specific prediction coding scheme is shown in formula (X).
  • the final best coding mode is passed to the decoding end, and the decoding end uses the obtained best coding mode to reconstruct the attributes of the point cloud, thereby further improving the coding efficiency of the point cloud attributes.
  • Table 3 shows the test results on the coding efficiency of the attributes.
  • the attribute coding BPP is reduced by about 3.9%, which significantly improves the coding efficiency of point cloud attributes.
  • a code stream is provided, wherein the code stream is generated by bit encoding according to information to be encoded; wherein the information to be encoded includes at least one of the following: a value of the first grammar identification information, a value of the second grammar identification information, a value of the third grammar identification information, a value of the fourth grammar identification information, a value of the fifth grammar identification information, a weight index value corresponding to a node in the current layer, and a second coefficient quantized residual value of the node in the current layer;
  • the first grammatical identification information is used to indicate whether the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode
  • the second grammatical identification information is used to indicate the target coding mode of the current layer
  • the value of the third grammatical identification information is used to indicate the number of layers included in the current sequence where the current layer is located
  • the value of the fourth grammatical identification information is used to indicate whether the nodes of the current layer are allowed to perform inter-frame prediction
  • the fifth grammatical identification information is used to indicate whether the nodes of the current layer are allowed to perform intra-frame prediction
  • the sixth grammatical identification information is used to indicate that the nodes in the current layer adopt the regional adaptive layered inter-frame transformation mode
  • the weight index value is used to indicate the index value corresponding to the target weight combination corresponding to the nodes in the current layer in the preset weight table.
  • FIG. 44 shows a schematic diagram of the composition structure of a decoder provided by the embodiment of the present application.
  • the decoder 1000 may include a first determining part 1001 and a decoding part 1002; wherein,
  • the first determining part 1001 is configured to parse the bitstream and determine the first syntax identification information when it is determined that the node of the current layer allows attribute prediction; and parse the bitstream and determine the target decoding mode of the current layer when the first syntax identification information indicates that the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode;
  • the decoding part 1002 is configured to perform attribute decoding on the nodes in the current layer according to the target decoding mode, and determine the attribute reconstruction values of the nodes in the current layer.
  • the first determination part 1001 is further configured to determine that the current coefficient group allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode if the value of the first syntax identification information is a first value; the current coefficient group includes at least one layer, and the current layer is one of the at least one layer; if the value of the first syntax identification information is a second value, determine that the current coefficient group does not allow adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode.
  • the first determination part 1001 is further configured to determine that the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode if the value of the first syntax identification information is a first value; and to determine that the current layer does not allow adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode if the value of the first syntax identification information is a second value.
  • the first determination part 1001 is further configured to decode the code stream, determine the attribute block header information parameter set; determine the second syntax identification information from the attribute block header information parameter set; and determine the target decoding mode of the current layer based on the second syntax identification information.
  • the target decoding mode includes a region adaptive hierarchical intra-frame transform mode, a region adaptive hierarchical inter-frame transform mode and a region adaptive hierarchical combined transform mode; wherein, the region adaptive hierarchical intra-frame transform mode represents the use of an intra-frame prediction mode to perform attribute prediction transform decoding on the nodes of the current layer; the region adaptive hierarchical inter-frame transform mode represents the use of an inter-frame prediction mode to perform attribute prediction transform decoding on the nodes of the current layer; the region adaptive hierarchical combined transform mode represents the use of an intra-frame prediction mode combined with an inter-frame prediction mode to perform attribute prediction transform decoding on the nodes of the current layer.
  • the region adaptive hierarchical inter-frame transform mode includes a first region adaptive hierarchical inter-frame transform mode and a second region adaptive hierarchical inter-frame transform mode;
  • the region adaptive hierarchical combined transform mode includes a first region adaptive hierarchical combined transform mode, a second region adaptive hierarchical combined transform mode and a third region adaptive hierarchical combined transform mode;
  • the first region adaptive hierarchical inter-frame transform mode represents the attribute prediction transform decoding of the nodes of the current layer in a manner of determining the co-located prediction nodes by using the geometric information of the nodes;
  • the second region adaptively hierarchically combines the transform mode characterization to perform attribute prediction transform decoding on the nodes of the current layer by using the cache of the reference frame to determine the co-located prediction nodes;
  • the first region adaptive hierarchical combined transform mode characterizes the use of a combined region adaptive hierarchical intra-frame transform mode and the first region adaptive hierarchical inter-frame transform mode to perform attribute prediction transform decoding on the nodes of the current layer;
  • the second region adaptive hierarchical combined transform mode characterizes the use of a combined region adaptive hierarchical intra-frame transform mode and the second region adaptive hierarchical inter-frame transform mode to perform attribute prediction transform decoding on the nodes of the current layer;
  • the third region adaptive hierarchical combined transform mode characterization uses a combined region adaptive hierarchical intra-frame transform mode, the first region adaptive hierarchical inter-frame transform mode and the second region adaptive hierarchical inter-frame transform mode to perform attribute prediction transform decoding on the nodes of the current layer.
  • the first determination part 1001 is further configured to determine the target decoding mode of the current layer using the inter-frame prediction mode and/or the intra-frame prediction mode according to the value of the second syntax identification information.
  • the first determining part 1001 is further configured to determine that the target decoding mode of the current layer is a region adaptive hierarchical intra transform mode if the value of the second syntax identification information is a third value;
  • the value of the second syntax identification information is a fourth value, determining that the target decoding mode of the current layer is a first region adaptive hierarchical inter-frame transform mode
  • the target decoding mode of the current layer is a second region adaptive hierarchical inter-frame transform mode
  • the value of the second syntax identification information is the sixth value, determining that the target decoding mode of the current layer is the first region adaptive hierarchical combined transform mode
  • the value of the second syntax identification information is the seventh value, determining that the target decoding mode of the current layer is a second region adaptive hierarchical combined transform mode
  • the target decoding mode of the current layer is the third region adaptive hierarchical combined transform mode.
  • the first determination part 1001 is further configured to parse the code stream to determine the value of the third syntax identification information; the third syntax identification information is used to indicate the number of layers included in the current sequence where the current layer is located; obtain the index value of the current layer, if the index value of the current layer is greater than or equal to a ninth value and less than the number of layers, execute the step of parsing the code stream to determine the target decoding mode of the current layer; if the index value of the current layer is greater than the number of layers, do not execute the step of parsing the code stream to determine the target decoding mode of the current layer.
  • the first determination part 1001 is also configured to determine the attribute prediction values of the nodes in the current layer; forward transform the attribute prediction values of the nodes in the current layer according to the target decoding mode to determine the first coefficient values and the second coefficient prediction values of the nodes in the current layer; determine the second coefficient values corresponding to the nodes in the current layer according to the second coefficient prediction values; inversely transform the first coefficient values and the second coefficient values of the nodes in the current layer according to the target decoding mode to determine the attribute reconstruction values of the nodes in the current layer.
  • the first determination part 1001 is also configured to determine the adjacent nodes of the nodes in the current layer; wherein the adjacent nodes include neighboring nodes and parent node neighboring nodes; linear fitting is performed based on the attribute reconstruction values corresponding to the adjacent nodes and the geometric distances between the nodes in the current layer and the adjacent nodes to determine the attribute prediction values of the nodes in the current layer.
  • the first determination part 1001 is further configured to decode the code stream to determine the second coefficient decoding residual value of the node in the current layer; perform inverse quantization on the second coefficient decoding residual value to obtain the second coefficient inverse quantization residual value of the node in the current layer; determine the second coefficient value of the node in the current layer based on the second coefficient prediction value corresponding to the node in the current layer and the second coefficient inverse quantization residual value.
  • the first determination part 1001 is also configured to perform a forward transform on the nodes in the current layer according to the region adaptive layered combined transform mode, and determine the first intermediate prediction value and the second intermediate prediction value of the nodes in the current layer; add the first intermediate prediction value and the second intermediate prediction value of the nodes in the current layer to obtain the second coefficient prediction value of the nodes in the current layer.
  • the first determination part 1001 is also configured to determine the first target weight and the second target weight of the current layer; adopt the region adaptive layered intra-frame transformation mode to forward transform the nodes in the current layer to determine the first attribute prediction value of the nodes in the current layer; adopt the region adaptive layered inter-frame transformation mode to forward transform the nodes in the current layer to determine the second attribute prediction value of the nodes in the current layer; multiply the first attribute prediction value of the nodes in the current layer by the first target weight to obtain the first intermediate prediction value of the nodes in the current layer; multiply the second attribute prediction value of the nodes in the current layer by the second target weight to obtain the second intermediate prediction value of the nodes in the current layer.
  • the first determination part 1001 is further configured to parse the code stream and determine the weight index value; determine the target weight combination corresponding to the weight index in the preset weight table; wherein the target weight combination includes a first target weight and a second target weight.
  • the first determination part 1001 is also configured to, when the second coefficient prediction value of the node in the current layer is the tenth value, adopt the region adaptive hierarchical intra-frame transform mode to perform a forward transform on the node in the current layer to obtain an intermediate second coefficient prediction value of the node in the current layer; and use the intermediate second coefficient prediction value as the second coefficient prediction value of the node in the current layer.
  • the first grammar identification information when the fourth grammar identification information and the fifth grammar identification information are both first values, the first grammar identification information is the first value; the fourth grammar identification information is used to indicate whether the nodes of the current layer are allowed to perform inter-frame prediction; the fifth grammar identification information is used to indicate whether the nodes of the current layer are allowed to perform intra-frame prediction; when either the fourth grammar identification information or the fifth grammar identification information is a second value, the first grammar identification information is a second value.
  • the first determining part 1001 is further configured to: When the current layer does not allow adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode, parse the code stream to determine fourth syntax identification information; when the fourth syntax identification information indicates that the nodes of the current layer do not allow inter-frame prediction, parse the code stream to determine fifth syntax identification information; when the fifth syntax identification information indicates that the nodes of the current layer allow intra-frame prediction, perform attribute decoding on the nodes in the current layer according to the region adaptive hierarchical intra-frame transform mode to determine the attribute reconstruction values of the nodes in the current layer.
  • the first determination part 1001 is also configured to perform attribute decoding on the nodes in the current layer according to the region adaptive layered inter-frame transformation mode to determine the attribute reconstruction values of the nodes in the current layer when the fourth syntax identification information indicates that the nodes in the current layer allow inter-frame prediction.
  • the first determination part 1001 is further configured to determine that the nodes of the current layer are allowed to perform inter-frame prediction if the value of the fourth grammar identification information is a first value; and to determine that the nodes of the current layer are not allowed to perform inter-frame prediction if the value of the fourth grammar identification information is a second value.
  • the first determination part 1001 is further configured to determine that the nodes of the current layer are allowed to perform inter-frame prediction if the value of the fifth grammar identification information is a first value; and to determine that the nodes of the current layer are not allowed to perform inter-frame prediction if the value of the fifth grammar identification information is a second value.
  • the first determination part 1001 is further configured to parse the code stream to determine the value of the sixth grammar identification information; the sixth grammar identification information is used to indicate that the nodes in the current layer adopt the region adaptive layered inter-frame transform mode.
  • the first determination part 1001 is also configured to determine the number of adjacent nodes of the current layer; wherein the adjacent nodes include the number of neighboring nodes and the number of parent node neighboring nodes; when the number of adjacent nodes is greater than or equal to a preset threshold, it is determined that the nodes of the current layer allow attribute prediction.
  • the first determination part 1001 is also configured to determine the neighboring nodes of each node based on the spatial position of each node in the current layer; wherein the neighboring nodes of the node include at least: neighboring nodes coplanar with the node and neighboring nodes colinear with the node.
  • the first determination part 1001 is also configured to determine the parent node of each of the nodes in the current layer; determine the parent node neighboring nodes of each of the nodes based on the spatial position of the parent node of each of the nodes; wherein the parent node neighboring nodes of the node include at least: neighboring nodes coplanar with the parent node of the node and neighboring nodes colinear with the parent node of the node.
  • the first determination part 1001 is also configured to count the number of neighboring nodes of each node in the current layer to determine the number of neighboring nodes of the current layer; count the number of neighboring nodes of the parent node of each node in the current layer to determine the number of neighboring nodes of the parent node of the current layer; add the number of neighboring nodes and the number of neighboring nodes of the parent node to obtain the number of adjacent nodes of the current layer.
  • part can be a part of a circuit, a part of a processor, a part of a program or software, etc., and of course it can also be a module, or it can be non-modular.
  • the components in the present embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
  • the above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional module.
  • the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium.
  • the technical solution of this embodiment is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product.
  • the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method described in this embodiment.
  • the aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
  • an embodiment of the present application provides a computer-readable storage medium, which is applied to the decoder 1000.
  • the computer-readable storage medium stores a computer program, and when the computer program is executed by the first processor, the method described in any one of the aforementioned embodiments is implemented.
  • the decoder 1000 may include: a first communication interface 1101, a first memory 1102 and a first processor 1103; each component is coupled together through a first bus system 1104. It can be understood that the first bus system 1104 is used to realize the connection and communication between these components.
  • the first bus system 1104 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are marked as the first bus system 1104 in Figure 11. Among them,
  • the first communication interface 1101 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;
  • a first memory 1102 used to store a computer program that can be run on the first processor 1103;
  • the first processor 1103 is configured to, when running the computer program, execute:
  • the bitstream is parsed to determine the first syntax identification information
  • the first syntax identification information indicates that the current layer allows adaptive selection of an inter-frame prediction mode and/or an intra-frame prediction mode, parsing a bitstream to determine a target decoding mode of the current layer;
  • Attribute decoding is performed on the nodes in the current layer according to the target decoding mode to determine attribute reconstruction values of the nodes in the current layer.
  • the first memory 1102 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.
  • the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.
  • the volatile memory can be a random access memory (RAM), which is used as an external cache.
  • RAM static RAM
  • DRAM dynamic RAM
  • SDRAM synchronous DRAM
  • DDRSDRAM double data rate synchronous DRAM
  • ESDRAM enhanced synchronous DRAM
  • SLDRAM synchronous link DRAM
  • DRRAM direct RAM bus RAM
  • the first processor 1103 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the first processor 1103.
  • the above-mentioned first processor 1103 can be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
  • DSP Digital Signal Processor
  • ASIC Application Specific Integrated Circuit
  • FPGA Field Programmable Gate Array
  • the methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed.
  • the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
  • the steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor can be executed.
  • the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc.
  • the storage medium is located in the first memory 1102, and the first processor 1103 reads the information in the first memory 1102 and completes the steps of the above method in combination with its hardware.
  • the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processors (Digital Signal Processing, DSP), digital signal processing devices (DSP Device, DSPD), programmable logic devices (Programmable Logic Device, PLD), field programmable gate arrays (Field-Programmable Gate Array, FPGA), general processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application or a combination thereof.
  • ASIC Application Specific Integrated Circuits
  • DSP Digital Signal Processing
  • DSP Device digital signal processing devices
  • PLD programmable logic devices
  • FPGA field programmable gate array
  • general processors controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application or a combination thereof.
  • the technology described in this application can be implemented by a module (such as a process, function, etc.) that performs the functions described in this application.
  • the software code can be stored in a memory and executed by a processor.
  • the memory can be implemented in the processor or outside the processor.
  • the first processor 1103 is further configured to execute any one of the methods described in the foregoing embodiments when running the computer program.
  • This embodiment provides a decoder, in which a corresponding attribute decoding mode is introduced for each layer.
  • the target decoding mode of each layer can be adaptively selected at the decoding end, so that the decoding end uses the parsed target decoding mode to reconstruct the attributes of the point cloud, thereby improving the decoding efficiency of the point cloud attributes, and further improving the decoding performance of the point cloud.
  • FIG46 shows a schematic diagram of the composition structure of an encoder provided by an embodiment of the present application.
  • the encoder 2000 may include a second determination part 2001 and an encoding part 2002; wherein,
  • the second determination part is used to determine the target coding mode of the current layer and determine the first syntax identification information when it is determined that the node of the current layer allows attribute prediction and the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode; the first syntax identification information is used to indicate whether the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode;
  • the encoding part is used to perform attribute encoding on the nodes in the current layer according to the target encoding mode, and determine the attribute reconstruction values of the nodes in the current layer.
  • the encoding part 2002 is configured to determine that the value of the first grammar identification information is a first value if it is determined that the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode; and to determine that the value of the first grammar identification information is a second value if it is determined that the current layer does not allow adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode.
  • the second determining part 2001 is further configured to determine that the current coefficient group allows adaptive selection. if the current coefficient group does not allow adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode, then the value of the first syntax identification information is determined to be a first value; the current coefficient group includes at least one layer, and the current layer is one of the at least one layer; if it is determined that the current coefficient group does not allow adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode, then the value of the first syntax identification information is determined to be a second value.
  • the second determination part 2001 is further configured to determine that the value of the first grammar identification information is a first value if it is determined that the current layer allows adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode; and to determine that the value of the first grammar identification information is a second value if it is determined that the current layer does not allow adaptive selection of the inter-frame prediction mode and/or the intra-frame prediction mode.
  • the second determination part 2001 is further configured to determine second syntax identification information according to the target coding mode; add the second syntax identification information to the attribute block header information parameter set, encode the attribute block header information parameter set, and write the obtained encoded bits into the bitstream.
  • the target coding mode includes a region adaptive hierarchical intra-frame transform mode, a region adaptive hierarchical inter-frame transform mode and a region adaptive hierarchical combined transform mode; wherein, the region adaptive hierarchical intra-frame transform mode represents the use of an intra-frame prediction mode to perform attribute prediction transform encoding on the nodes of the current layer; the region adaptive hierarchical inter-frame transform mode represents the use of an inter-frame prediction mode to perform attribute prediction transform encoding on the nodes of the current layer; the region adaptive hierarchical combined transform mode represents the use of an intra-frame prediction mode combined with an inter-frame prediction mode to perform attribute prediction transform encoding on the nodes of the current layer.
  • the region adaptive hierarchical inter-frame transform mode includes a first region adaptive hierarchical inter-frame transform mode and a second region adaptive hierarchical inter-frame transform mode;
  • the region adaptive hierarchical combined transform mode includes a first region adaptive hierarchical combined transform mode, a second region adaptive hierarchical combined transform mode and a third region adaptive hierarchical combined transform mode;
  • the first region adaptive hierarchical inter-frame transform mode represents the attribute prediction transform coding of the nodes of the current layer in a manner of determining the co-located prediction nodes by using the geometric information of the nodes;
  • the second region adaptively hierarchically combines the transform mode characterization to perform attribute prediction transform coding on the nodes of the current layer by using the cache of the reference frame to determine the co-located prediction nodes;
  • the first region adaptive hierarchical combined transform mode characterizes that attribute prediction transform coding is performed on nodes of the current layer by combining the region adaptive hierarchical intra transform mode and the first region adaptive hierarchical inter transform mode;
  • the second region adaptive hierarchical combined transform mode characterizes that attribute prediction transform coding is performed on the nodes of the current layer by combining the region adaptive hierarchical intra transform mode and the second region adaptive hierarchical inter transform mode;
  • the third region adaptive hierarchical combined transform mode characterization uses a combined region adaptive hierarchical intra-frame transform mode, the first region adaptive hierarchical inter-frame transform mode and the second region adaptive hierarchical inter-frame transform mode to perform attribute prediction transform coding on the nodes of the current layer.
  • the second determination part 2001 is further configured to determine the value of the second syntax identification information according to the target decoding mode of the current layer adopting the inter-frame prediction mode and/or the intra-frame prediction mode.
  • the second determining part 2001 is further configured to determine that the value of the second syntax identification information is a third value if the target coding mode of the current layer is a region adaptive hierarchical intra transform mode;
  • the target coding mode of the current layer is a first region adaptive hierarchical inter-frame transform mode, determining that the value of the second syntax identification information is a fourth value;
  • the target coding mode of the current layer is a second region adaptive hierarchical inter-frame transform mode, determining that the value of the second syntax identification information is a fifth value;
  • the target coding mode of the current layer is the first region adaptive hierarchical combined transform mode, determining that the value of the second syntax identification information is a sixth value;
  • the target coding mode of the current layer is the second region adaptive hierarchical combined transform mode, determining that the value of the second syntax identification information is a seventh value
  • the value of the second syntax identification information is determined to be the eighth value.
  • the second determination part 2001 is further configured to determine the value of third grammatical identification information; the third grammatical identification information is used to indicate the number of layers included in the current sequence where the current layer is located; obtain the index value of the current layer, if the index value of the current layer is greater than or equal to the ninth value and less than the number of layers, execute the step of determining the target coding mode of the current layer; if the index value of the current layer is greater than the number of layers, do not execute the step of determining the target coding mode of the current layer; encode the third grammatical identification information, and write the obtained coded bits into the bitstream.
  • the second determination part 2001 is further configured to perform attribute encoding on the current layer based on at least one candidate coding mode, and determine a precoding result of each of the at least one candidate coding mode; perform cost calculation according to the precoding result of each of the at least one candidate coding mode, and determine a cost value of each of the at least one candidate coding mode; and determine the target coding mode of the current layer from the at least one candidate coding mode according to the cost value of each of the at least one candidate coding mode.
  • the second determining part 2001 is further configured to determine the encoding mode according to the at least one candidate encoding mode.
  • the respective precoding results are subjected to rate-distortion cost calculation to determine the respective cost value of the at least one candidate coding mode.
  • the second determination part 2001 is further configured to determine a minimum cost value from the cost values of the at least one candidate coding mode; and determine the candidate coding mode corresponding to the minimum cost value as the target coding mode of the current layer.
  • the at least one candidate encoding mode includes at least one of the following: a region adaptive layered intra-frame transform mode, a first region adaptive layered inter-frame transform mode, a second region adaptive layered inter-frame transform mode, a first region adaptive layered combined transform mode, a second region adaptive layered combined transform mode, and a third region adaptive layered combined transform mode.
  • the encoding part 2002 is also configured to determine the attribute prediction value of the node in the current layer; forward transform the attribute prediction value of the node in the current layer according to the target encoding mode to determine the first coefficient value and the second coefficient prediction value of the node in the current layer; determine the second coefficient value corresponding to the node in the current layer according to the second coefficient prediction value; inversely transform the first coefficient value and the second coefficient value of the node in the current layer according to the target encoding mode to determine the attribute reconstruction value of the node in the current layer.
  • the encoding part 2002 is also configured to determine the adjacent nodes of the nodes in the current layer; wherein the adjacent nodes include neighboring nodes and parent node neighboring nodes; linear fitting is performed based on the attribute reconstruction values corresponding to the adjacent nodes and the geometric distance between the nodes in the current layer and the adjacent nodes to determine the attribute prediction values of the nodes in the current layer.
  • the encoding part 2002 is also configured to determine the second coefficient encoding residual value of the node in the current layer; perform inverse quantization on the second coefficient encoding residual value to obtain the second coefficient inverse quantization residual value of the node in the current layer; and determine the second coefficient value of the node in the current layer based on the second coefficient prediction value corresponding to the node in the current layer and the second coefficient inverse quantization residual value.
  • the encoding part 2002 is also configured to determine the original value of the attribute of the node in the current layer; forward transform the original value of the attribute of the node in the current layer according to the target coding mode to determine the first coefficient value and the second coefficient original value of the node in the current layer; determine the second coefficient prediction residual value of the node in the current layer according to the second coefficient original value and the second coefficient prediction value of the point in the current layer; quantize the second coefficient residual value to obtain the second coefficient quantized residual value of the node in the current layer.
  • the encoding part 2002 is further configured to encode the second coefficient quantization residual value of the node in the current layer, and write the obtained encoding bits into the bit stream.
  • the encoding part 2002 when the target coding mode of the current layer is the region adaptive layered combined transform mode, is also configured to perform a forward transform on the nodes in the current layer according to the region adaptive layered combined transform mode, determine the first intermediate prediction value and the second intermediate prediction value of the nodes in the current layer; add the first intermediate prediction value and the second intermediate prediction value of the nodes in the current layer to obtain the second coefficient prediction value of the nodes in the current layer.
  • the encoding part 2002 is also configured to determine a target weight combination corresponding to the current layer in a preset weight table; wherein the target weight combination includes a first target weight and a second target weight; using the regional adaptive layered intra-frame transform mode, forward transforming the nodes in the current layer to determine the first attribute prediction value of the nodes in the current layer; using the regional adaptive layered inter-frame transform mode, forward transforming the nodes in the current layer to determine the second attribute prediction value of the nodes in the current layer; multiplying the first attribute prediction value of the nodes in the current layer by the first target weight to obtain the first intermediate prediction value of the nodes in the current layer; multiplying the second attribute prediction value of the nodes in the current layer by the second target weight to obtain the second intermediate prediction value of the nodes in the current layer.
  • the encoding part 2002 is further configured to encode the target weight combination with the corresponding weight index value in the preset weight table, and write the obtained encoded bits into the bitstream.
  • the encoding part 2002 is also configured to determine the attribute prediction value of the node in the current layer; forward transform the attribute prediction value of the node in the current layer according to the target encoding mode to determine the first coefficient value and the second coefficient prediction value of the node in the current layer; determine the second coefficient value corresponding to the node in the current layer according to the second coefficient prediction value; inversely transform the first coefficient value and the second coefficient value of the node in the current layer according to the target encoding mode to determine the attribute reconstruction value of the node in the current layer.
  • the encoding part 2002 is also configured to determine the adjacent nodes of the nodes in the current layer; wherein the adjacent nodes include neighboring nodes and parent node neighboring nodes; linear fitting is performed based on the attribute reconstruction values corresponding to the adjacent nodes and the geometric distance between the nodes in the current layer and the adjacent nodes to determine the attribute prediction values of the nodes in the current layer.
  • the encoding part 2002 is also configured to determine the second coefficient encoding residual value of the node in the current layer; perform inverse quantization on the second coefficient encoding residual value to obtain the second coefficient inverse quantization residual value of the node in the current layer; and determine the second coefficient value of the node in the current layer based on the second coefficient prediction value corresponding to the node in the current layer and the second coefficient inverse quantization residual value.
  • the encoding part 2002 is further configured to determine the attribute original of the node in the current layer. value; forward transform the original value of the attribute of the node in the current layer according to the target coding mode, and determine the first coefficient value and the second coefficient original value of the node in the current layer; determine the second coefficient prediction residual value of the node in the current layer according to the second coefficient original value and the second coefficient prediction value of the point in the current layer; quantize the second coefficient residual value to obtain the second coefficient quantized residual value of the node in the current layer.
  • the encoding part 2002 is further configured to encode the second coefficient quantization residual value of the node in the current layer, and write the obtained encoding bits into the bit stream.
  • the encoding part 2002 when the target coding mode of the current layer is the region adaptive layered combined transform mode, is also configured to perform a forward transform on the nodes in the current layer according to the region adaptive layered combined transform mode, determine the first intermediate prediction value and the second intermediate prediction value of the nodes in the current layer; add the first intermediate prediction value and the second intermediate prediction value of the nodes in the current layer to obtain the second coefficient prediction value of the nodes in the current layer.
  • the encoding part 2002 is also configured to determine a target weight combination corresponding to the current layer in a preset weight table; wherein the target weight combination includes a first target weight and a second target weight; using the regional adaptive layered intra-frame transform mode, forward transforming the nodes in the current layer to determine the first attribute prediction value of the nodes in the current layer; using the regional adaptive layered inter-frame transform mode, forward transforming the nodes in the current layer to determine the second attribute prediction value of the nodes in the current layer; multiplying the first attribute prediction value of the nodes in the current layer by the first target weight to obtain the first intermediate prediction value of the nodes in the current layer; multiplying the second attribute prediction value of the nodes in the current layer by the second target weight to obtain the second intermediate prediction value of the nodes in the current layer.
  • the encoding part 2002 is further configured to encode the target weight combination with the corresponding weight index value in the preset weight table, and write the obtained encoded bits into the bitstream.
  • the encoding part 2002 when the target coding mode of the current layer is the region adaptive hierarchical inter-frame transform mode or the region adaptive hierarchical combined transform mode, is also configured to, when the second coefficient prediction value of the node in the current layer is the tenth value, perform a forward transform on the node in the current layer according to the region adaptive hierarchical intra-frame transform mode to obtain an intermediate second coefficient prediction value of the node in the current layer; and use the intermediate second coefficient prediction value as the second coefficient prediction value of the node in the current layer.
  • the encoding part 2002 is also configured to determine the values of fourth grammar identification information and fifth grammar identification information; the fifth grammar identification information is used to indicate whether the node of the current layer is allowed to perform intra-frame prediction; when the fourth grammar identification information and the fifth grammar identification information are both first values, the value of the first grammar identification information is determined to be the first value; when either the fourth grammar identification information or the fifth grammar identification information is the second value, the value of the first grammar identification information is determined to be the second value; the first grammar identification information is encoded and the obtained encoded bits are written into the bitstream.
  • the second determination part 2001 is further configured to set the value of the fourth grammar identification information to the first value if it is determined that the nodes of the current layer allow inter-frame prediction; set the value of the fourth grammar identification information to the second value if it is determined that the nodes of the current layer do not allow inter-frame prediction; encode the fourth grammar identification information and write the obtained coded bits into the bitstream.
  • the second determination part 2001 is further configured to set the value of the fifth grammar identification information to the first value if it is determined that the nodes of the current layer allow intra-frame prediction; set the value of the fifth grammar identification information to the second value if it is determined that the nodes of the current layer do not allow intra-frame prediction; encode the fifth grammar identification information and write the obtained coded bits into the bitstream.
  • the second determination part 2001 is further configured to determine the value of sixth grammar identification information; the sixth grammar identification information is used to indicate that the nodes in the current layer adopt the regional adaptive layered inter-frame transformation mode; the sixth grammar identification information is encoded and the obtained encoded bits are written into the bitstream.
  • the second determining part 2001 is further configured to perform encoding processing on the target encoding mode of the node in the current layer, and write the obtained encoding bits into the bitstream.
  • the second determination part 2001 is further configured to determine the number of adjacent nodes of the current layer; wherein the adjacent nodes include the number of neighboring nodes and the number of parent node neighboring nodes; when the number of adjacent nodes is greater than or equal to a preset threshold, it is determined that the nodes of the current layer allow attribute prediction.
  • the second determination part 2001 is also configured to determine the neighboring nodes of each node based on the spatial position of each node in the current layer; wherein the neighboring nodes of the node include at least: neighboring nodes coplanar with the node and neighboring nodes colinear with the node.
  • the second determination part 2001 is also configured to determine the parent node of each of the nodes in the current layer; determine the parent node neighboring nodes of each of the nodes based on the spatial position of the parent node of each of the nodes; wherein the parent node neighboring nodes of the node include at least: neighboring nodes coplanar with the parent node of the node and neighboring nodes colinear with the parent node of the node.
  • the second determination part 2001 is also configured to count the number of neighboring nodes of each node in the current layer to determine the number of neighboring nodes of the current layer; count the number of neighboring nodes of the parent node of each node in the current layer to determine the number of neighboring nodes of the parent node of the current layer; add the number of neighboring nodes and the number of neighboring nodes of the parent node to obtain the number of adjacent nodes of the current layer.
  • part can be a part of the circuit, a part of the processor, a part of the program or software, etc., and of course it can also be a module, or it can be non-modular.
  • the components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
  • the above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional modules.
  • the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium.
  • this embodiment provides a computer-readable storage medium, which is applied to the encoder 2000, and the computer-readable storage medium stores a computer program. When the computer program is executed by the second processor, the method described in any one of the above embodiments is implemented.
  • the encoder 2000 may include: a second communication interface 2101, a second memory 2102 and a second processor 2103; each component is coupled together through a second bus system 2104. It can be understood that the second bus system 2104 is used to realize the connection and communication between these components.
  • the second bus system 2104 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are marked as the second bus system 2104 in Figure 21. Among them,
  • the second communication interface 2101 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;
  • the second memory 2102 is used to store a computer program that can be run on the second processor 2103;
  • the second processor 2103 is configured to, when running the computer program, execute:
  • a target coding mode of the current layer is determined, and first syntax identification information is determined; the first syntax identification information is used to indicate whether the current layer allows adaptive selection of an inter-frame prediction mode and/or an intra-frame prediction mode;
  • Attribute encoding is performed on the nodes in the current layer according to the target encoding mode to determine attribute reconstruction values of the nodes in the current layer.
  • the second processor 2103 is further configured to execute any one of the methods described in the foregoing embodiments when running the computer program.
  • This embodiment provides an encoder, in which a corresponding attribute coding mode is introduced for each layer.
  • the encoding end can adaptively select the target coding mode of each layer, so that the encoding end uses the parsed target coding mode to reconstruct the attributes of the point cloud, thereby improving the coding efficiency of the point cloud attributes, and further improving the coding performance of the point cloud.
  • FIG48 shows a schematic diagram of the composition structure of a coding and decoding system provided in an embodiment of the present application.
  • the coding and decoding system 3000 may include a decoder 3001 and an encoder 3002 .
  • the decoder 3001 may be the decoder described in any one of the aforementioned embodiments
  • the encoder 3002 may be the encoder described in any one of the aforementioned embodiments.
  • the code stream when it is determined that the nodes of the current layer allow attribute prediction, the code stream is parsed to determine the first syntax identification information; when the first syntax identification information indicates that the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode, the code stream is parsed to determine the target decoding mode of the current layer; the nodes in the current layer are attribute-decoded according to the target decoding mode to determine the attribute reconstruction values of the nodes in the current layer.
  • the target encoding mode of the current layer is determined, and the first syntax identification information is determined; the first syntax identification information is used to indicate whether the current layer allows adaptive selection of inter-frame prediction mode and/or intra-frame prediction mode; the nodes in the current layer are attribute-encoded according to the target encoding mode to determine the attribute reconstruction values of the nodes in the current layer.
  • the encoding end when performing attribute encoding for each layer, can adaptively select the target coding mode of each slice, and pass the target coding mode to the decoding end, so that the decoding end uses the parsed target decoding mode to reconstruct the attributes of the point cloud, thereby improving the encoding and decoding efficiency of the point cloud attributes, and then improving the encoding and decoding performance of the point cloud.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computing Systems (AREA)
  • Theoretical Computer Science (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

本申请实施例公开了一种编解码方法、码流、编码器、解码器以及存储介质,该方法包括:在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;在第一语法标识信息指示当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定当前层的目标解码模式;根据目标解码模式对当前层中的节点进行属性解码,确定当前层中的节点的属性重建值。这样,可以提升点云属性的编解码效率,进而提升点云的编解码性能。

Description

编解码方法、码流、编码器、解码器以及存储介质 技术领域
本申请实施例涉及点云编解码技术领域,尤其涉及一种编解码方法、码流、编码器、解码器以及存储介质。
背景技术
在运动图像专家组(Moving Picture Experts Group,MPEG)提供的基于几何的点云压缩(Geometry-based Point Cloud Compression,G-PCC)编解码框架或基于视频的点云压缩(Video-based Point Cloud Compression,V-PCC)编解码框架中,点云的几何信息和属性信息是分开进行编码的。
目前,属性信息编码主要针对颜色信息的编码,在颜色信息编码中,主要有两种变换方法,一是依赖于细节层次(Level of Detail,LOD)划分的基于距离的提升变换,另一是直接进行的区域自适应分层变换(Region Adaptive Hierarchal Transform,RAHT)。
然而,在进行RAHT变换时,可以通过属性参数集(Attribute Parameters Set,APS)来确定整个序列的属性编码方式,例如,整个序列是采用RATH变换,还是预测变换或者提升变换进行属性编码。但是这种属性编码方案并没有完全考虑到不同RAHT层的交流分量(Alternating Crrent)的分布情况,从而导致点云属性的编码效率较低。
发明内容
本申请实施例提供一种编解码方法、码流、编码器、解码器以及存储介质,可以提升点云属性的编码效率,进而提升点云的编解码性能。
本申请实施例的技术方案可以如下实现:
第一方面,本申请实施例提供了一种解码方法,应用于解码器,该方法包括:
在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;
在所述第一语法标识信息指示所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定所述当前层的目标解码模式;
根据所述目标解码模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值。
第二方面,本申请实施例提供了一种编码方法,应用于编码器,该方法包括:
在确定当前层的节点允许进行属性预测,以及所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,确定所述当前层的目标编码模式;
根据所述目标编码模式对所述当前层中的节点进行属性编码,确定所述当前层中的节点的属性重建值。
第三方面,本申请实施例提供了一种码流,该码流是根据待编码信息进行比特编码生成的;其中,待编码信息包括下述至少一项:第一语法标识信息的取值、第二语法标识信息的取值、第三语法标识信息的取值、第四语法标识信息的取值、第五语法标识信息的取值、所述当前层中的节点对应的权重索引值、所述当前层中的节点的第二系数量化残差值;
其中,所述第一语法标识信息用于指示所述当前层是否允许自适应选择帧间预测模式和/或帧内预测模式,所述第二语法标识信息用于指示当前层的目标编码模式,所述第三语法标识信息的取值用于指示所述当前层所在的当前序列中所包含层的数量,所述第四语法标识信息的取值用于指示所述当前层的节点是否允许进行帧间预测,所述第五语法标识信息用于指示所述当前层的节点是否允许进行帧内预测,所述第六语法标识信息用于指示所述当前层中的节点采用区域自适应分层帧间变换模式,所述权重索引值用于指示所述当前层中的节点对应的目标权重组合在预设权重表中对应的索引值。
第四方面,本申请实施例提供了一种解码器,该解码器包括第一确定单元和解码单元;其中,
所述第一确定部分,用于在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;在所述第一语法标识信息指示所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定所述当前层的目标解码模式;
所述解码部分,用于根据所述目标解码模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值。
第五方面,本申请实施例提供了一种解码器,该解码器包括第一存储器和第一处理器;其中,
第一存储器,用于存储能够在第一处理器上运行的计算机程序;
第一处理器,用于在运行计算机程序时,执行如第一方面所述的方法。
第六方面,本申请实施例提供了一种编码器,该编码器包括第二确定单元和编码单元;其中,
所述第二确定部分,用于在确定当前层的节点允许进行属性预测,以及所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,确定所述当前层的目标编码模式,并确定第一语法标识信息;所述第一语法标识信息用于指示所述当前层是否允许自适应选择帧间预测模式和/或帧内预测模式的情况;
所述编码部分,用于根据所述目标编码模式对所述当前层中的节点进行属性编码,确定所述当前层中的节点的属性重建值。
第七方面,本申请实施例提供了一种编码器,该编码器包括第二存储器和第二处理器;其中,
第二存储器,用于存储能够在第二处理器上运行的计算机程序;
第二处理器,用于在运行计算机程序时,执行如第二方面所述的方法。
第八方面,本申请实施例提供了一种计算机可读存储介质,该计算机可读存储介质存储有计算机程序,所述计算机程序被执行时实现如第一方面所述的方法、或者实现如第二方面所述的方法。
本申请实施例提供了一种编解码方法、码流、编码器、解码器以及存储介质,在解码端,在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;在第一语法标识信息指示当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定当前层的目标解码模式;根据目标解码模式对当前层中的节点进行属性解码,确定当前层中的节点的属性重建值。在编码端,在确定当前层的节点允许进行属性预测,以及当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,确定当前层的目标编码模式,并确定第一语法标识信息;第一语法标识信息用于指示当前层是否允许自适应选择帧间预测模式和/或帧内预测模式的情况;编码部分,用于根据目标编码模式对当前层中的节点进行属性编码,确定当前层中的节点的属性重建值。这样,通过为每一层引入对应的属性编码模式,针对每一个层进行属性编码时,在编码端可以自适应地选择每一个片的目标编码模式,并将目标编码模式传递给解码端,使得解码端使用所解析到的目标解码模式来对点云的属性进行属性重建,从而提升了点云属性的编解码效率,进而提升了点云的编解码性能。
附图说明
图1A为一种三维点云图像示意图;
图1B为一种三维点云图像的局部放大图;
图2A为一种点云图像的六个观看角度示意图;
图2B为一种点云图像对应的数据存储格式示意图;
图3为一种点云编解码的网络架构示意图;
图4A为一种G-PCC编码器的组成框架示意图;
图4B为一种G-PCC解码器的组成框架示意图;
图5A为一种Z轴方向的低平面位置示意图;
图5B为一种Z轴方向的高平面位置示意图;
图6为一种节点编码顺序示意图;
图7A为一种平面标识信息示意图;
图7B为另一种平面标识信息示意图;
图8为一种当前节点的兄弟姐妹节点示意图;
图9为一种激光雷达与节点的相交示意图;
图10为一种处于相同划分深度以及相同坐标的邻域节点示意图;
图11为一种当前节点位于父节点的低平面位置示意图;
图12为一种当前节点位于父节点的高平面位置示意图;
图13为一种激光雷达点云平面位置信息的预测编码示意图;
图14为一种IDCM编码示意图;
图15为一种旋转激光雷达获取点云的坐标转换示意图;
图16为一种X轴或Y轴方向的预测编码示意图;
图17A为一种通过水平方位角来进行预测Y平面的角度示意图;
图17B为一种通过水平方位角来进行预测X平面的角度示意图;
图18为另一种X轴或Y轴方向的预测编码示意图;
图19A为一种子块包括的三个交点示意图;
图19B为一种利用三个交点拟合的三角面片集示意图;
图19C为一种三角面片集的上采样示意图;
图20为一种基于距离的LOD构造过程的示意图;
图21为一种LOD生成过程的可视化结果示意图;
图22为一种属性预测的编码流程示意图;
图23为一种金字塔结构的组成示意图;
图24为另一种金字塔结构的组成示意图;
图25为一种层间最近邻查找的LOD结构示意图;
图26为一种基于空间关系进行最近邻查找结构示意图;
图27A为一种共面的空间关系示意图;
图27B为一种共面和共线的空间关系示意图;
图27C为一种共面、共线和共点的空间关系示意图;
图28为一种基于快速查找的层间预测示意图;
图29为一种属性层内最近邻查找的LOD结构示意图;
图30为一种基于快速查找的层内预测示意图;
图31为一种基于块进行邻域查找结构示意图;
图32为一种提升变换的编码流程示意图;
图33为一种RAHT变换结构示意图;
图34为一种RAHT沿x、y、z三方向的变换过程示意图;
图35A为一种RAHT正变换的过程示意图;
图35B为一种RAHT逆变换的过程示意图;
图36为本申请实施例提供的一种解码方法的流程示意图;
图37为一种属性编码块的结构示意图;
图38为一种RAHT属性预测变换编码的整体流程示意图;
图39为一种当前块的邻域预测关系示意图;
图40为一种属性变换系数的计算过程示意图;
图41为一种RAHT属性帧间预测编码的结构示意图;
图42为为本申请实施例提供的一种编码方法的流程示意图;
图43为一种属性编码层的示意图;
图44为本申请实施例提供的一种解码器的组成结构示意图;
图45为本申请实施例提供的一种解码器的具体硬件结构示意图;
图46为本申请实施例提供的一种编码器的组成结构示意图;
图47为本申请实施例提供的一种编码器的具体硬件结构示意图;
图48为本申请实施例提供的一种编解码系统的组成结构示意图。
具体实施方式
为了能够更加详尽地了解本申请实施例的特点与技术内容,下面结合附图对本申请实施例的实现进行详细阐述,所附附图仅供参考说明之用,并非用来限定本申请实施例。
除非另有定义,本文所使用的所有的技术和科学术语与属于本申请的技术领域的技术人员通常理解的含义相同。本文中所使用的术语只是为了描述本申请实施例的目的,不是旨在限制本申请。
在以下的描述中,涉及到“一些实施例”,其描述了所有可能实施例的子集,但是可以理解,“一些实施例”可以是所有可能实施例的相同子集或不同子集,并且可以在不冲突的情况下相互结合。
还需要指出,本申请实施例所涉及的术语“第一\第二\第三”仅是用于区别类似的对象,不代表针对对象的特定排序,可以理解地,“第一\第二\第三”在允许的情况下可以互换特定的顺序或先后次序,以使这里描述的本申请实施例能够以除了在这里图示或描述的以外的顺序实施。
点云(Point Cloud)是物体表面的三维表现形式,通过光电雷达、激光雷达、激光扫描仪、多视角相机等采集设备,可以采集得到物体表面的点云(数据)。
点云是空间中一组无规则分布的、表达三维物体或场景的空间结构及表面属性的离散点集,图1A展示了三维点云图像和图1B展示了三维点云图像的局部放大图,可以看到点云表面是由分布稠密的点所组成的。
二维图像在每一个像素点均有信息表达,分布规则,因此不需要额外记录其位置信息;然而点云中的点在三维空间中的分布具有随机性和不规则性,因此需要记录每一个点在空间中的位置,才能完整地表达一幅点云。与二维图像类似,采集过程中每一个位置均有对应的属性信息,通常为RGB颜色值,颜色值反映物体的色彩;对于点云来说,每一个点所对应的属性信息除了颜色信息以外,还有比较常见的是反射率(reflectance)值,反射率值反映物体的表面材质。因此,点云数据通常包括点的位置信息和点的属性信息。其中,点的位置信息也可称为点的几何信息。例如,点的几何信息可以是点的三维坐标信息(x,y,z)。点的属性信息可以包括颜色信息和/或反射率等等。例如,反射率可以是一维反射率信息(r);颜色信息可以是任意一种色彩空间上的信息,或者颜色信息也可以是三维颜色信息,如RGB信息。在这里,R表示红色(Red,R),G表示绿色(Green,G),B表示蓝色(Blue,B)。再如,颜色信息可以是亮度色度(YCbCr,YUV)信息。其中,Y表示明亮度(Luma),Cb(U)表示蓝色色差,Cr(V)表示红色色差。
根据激光测量原理得到的点云,点云中的点可以包括点的三维坐标信息和点的反射率值。再如,根据摄影测量原理得到的点云,点云中的点可以可包括点的三维坐标信息和点的三维颜色信息。再如,结合激光测量和摄影测量原理得到点云,点云中的点可以可包括点的三维坐标信息、点的反射率值和点的三维颜色信息。
如图2A和图2B所示为一幅点云图像及其对应的数据存储格式。其中,图2A提供了点云图像的六个观看角度,图2B由文件头信息部分和数据部分组成,头信息包含了数据格式、数据表示类型、点云总点数、以及点云所表示的内容。例如,点云为“.ply”格式,由ASCII码表示,总点数为207242,每个点具有三维坐标信息(x,y,z)和三维颜色信息(r,g,b)。
点云可以按获取的途径分为:
静态点云:即物体是静止的,获取点云的设备也是静止的;
动态点云:物体是运动的,但获取点云的设备是静止的;
动态获取点云:获取点云的设备是运动的。
例如,按点云的用途分为两大类:
类别一:机器感知点云,其可以用于自主导航系统、实时巡检系统、地理信息系统、视觉分拣机器人、抢险救灾机器人等场景;
类别二:人眼感知点云,其可以用于数字文化遗产、自由视点广播、三维沉浸通信、三维沉浸交互等点云应用场景。
点云可以灵活方便地表达三维物体或场景的空间结构及表面属性,并且由于点云通过直接对真实物体采样获得,在保证精度的前提下能提供极强的真实感,因而应用广泛,其范围包括虚拟现实游戏、计算机辅助设计、地理信息系统、自动导航系统、数字文化遗产、自由视点广播、三维沉浸远程呈现、生物组织器官三维重建等。
点云的采集主要有以下途径:计算机生成、3D激光扫描、3D摄影测量等。计算机可以生成虚拟三维物体及场景的点云;3D激光扫描可以获得静态现实世界三维物体或场景的点云,每秒可以获取百万级点云;3D摄影测量可以获得动态现实世界三维物体或场景的点云,每秒可以获取千万级点云。这些技术降低了点云数据获取成本和时间周期,提高了数据的精度。点云数据获取方式的变革,使大量点云数据的获取成为可能,伴随着应用需求的增长,海量3D点云数据的处理遭遇存储空间和传输带宽限制的瓶颈。
示例性地,以帧率为30帧每秒(fps)的点云视频为例,每帧点云的点数为70万,每个点具有坐标信息xyz(float)和颜色信息RGB(uchar),则10s点云视频的数据量大约为0.7million×(4Byte×3+1Byte×3)×30fps×10s=3.15GB,其中,1Byte为10bit;而YUV采样格式为4:2:0,帧率为24fps的1280×720二维视频,其10s的数据量约为1280×720×12bit×24fps×10s≈0.33GB,10s的两视角三维视频的数据量约为0.33×2=0.66GB。由此可见,点云视频的数据量远超过相同时长的二维视频和三维视频的数据量。因此,为更好地实现数据管理,节省服务器存储空间,降低服务器与客户端之间的传输流量及传输时间,点云压缩成为促进点云产业发展的关键问题。
也就是说,由于点云是海量点的集合,存储点云不仅会消耗大量的内存,而且不利于传输,也没有这么大的带宽可以支持将点云不经过压缩直接在网络层进行传输,因此,需要对点云进行压缩。
目前,可对点云进行压缩的点云编码框架可以是运动图像专家组(Moving Picture Experts Group,MPEG)提供的基于几何的点云压缩(Geometry-based Point Cloud Compression,G-PCC)编解码框架或 基于视频的点云压缩(Video-based Point Cloud Compression,V-PCC)编解码框架,也可以是AVS提供的AVS-PCC编解码框架。G-PCC编解码框架可用于针对第一类静态点云和第三类动态获取点云进行压缩,其可以是基于点云压缩测试平台(Test Model Compression 13,TMC13),V-PCC编解码框架可用于针对第二类动态点云进行压缩,其可以是基于点云压缩测试平台(Test Model Compression 2,TMC2)。故G-PCC编解码框架也称为点云编解码器TMC13,V-PCC编解码框架也称为点云编解码器TMC2。
本申请实施例提供了一种包含解码方法和编码方法的点云编解码系统的网络架构,图3为本申请实施例提供的一种点云编解码的网络架构示意图。如图3所示,该网络架构包括一个或多个电子设备13至1N和通信网络01,其中,电子设备13至1N可以通过通信网络01进行视频交互。电子设备在实施的过程中可以为各种类型的具有点云编解码功能的设备,例如,所述电子设备可以包括手机、平板电脑、个人计算机、个人数字助理、导航仪、数字电话、视频电话、电视机、传感设备、服务器等,本申请实施例不作限制。其中,本申请实施例中的解码器或编码器就可以为上述电子设备。
其中,本申请实施例中的电子设备具有点云编解码功能,一般包括点云编码器(即编码器)和点云解码器(即解码器)。
下面以G-PCC编解码框架为例进行相关技术的说明。
可以理解,在点云G-PCC编解码框架中,针对待编码的点云数据,首先通过片(slice)划分,将点云数据划分为多个slice。在每一个slice中,点云的几何信息和每个点所对应的属性信息是分开进行编码的。
图4A示出了一种G-PCC编码器的组成框架示意图。如图4A所示,在几何编码过程中,对几何信息进行坐标转换,使点云全都包含在一个包围盒(Bounding Box)中,然后再进行量化,这一步量化主要起到缩放的作用,由于量化取整,使得一部分点云的几何信息相同,于是再基于参数来决定是否移除重复点,量化和移除重复点这一过程又被称为体素化过程。接着对包围盒进行八叉树划分或者预测树构建。在该过程中,针对划分的叶子结点中的点进行算术编码,生成二进制的几何比特流;或者,针对划分产生的交点(Vertex)进行算术编码(基于交点进行表面拟合),生成二进制的几何比特流。在属性编码过程中,几何编码完成,对几何信息进行重建后,需要先进行颜色转换,将颜色信息(即属性信息)从RGB颜色空间转换到YUV颜色空间。然后,利用重建的几何信息对点云重新着色,使得未编码的属性信息与重建的几何信息对应起来。属性编码主要针对颜色信息进行,在颜色信息编码过程中,主要有两种变换方法,一是依赖于细节层次(Level of Detail,LOD)划分的基于距离的提升变换,二是直接进行区域自适应分层变换(Region Adaptive Hierarchal Transform,RAHT),这两种方法都会将颜色信息从空间域转换到频域,通过变换得到高频系数和低频系数,最后对系数进行量化,再对量化系数进行算术编码,可以生成二进制的属性比特流。
图4B示出了一种G-PCC解码器的组成框架示意图。如图4B所示,针对所获取的二进制比特流,首先对二进制比特流中的几何比特流和属性比特流分别进行独立解码。在对几何比特流的解码时,通过算术解码-重构八叉树/重构预测树-重建几何-坐标逆转换,得到点云的几何信息;在对属性比特流的解码时,通过算术解码-反量化-LOD划分/RAHT-颜色逆转换,得到点云的属性信息,基于几何信息和属性信息还原待编码的点云数据(即输出点云)。
需要说明的是,在如图4A或图4B所示,目前G-PCC的几何编解码可以分为基于八叉树的几何编解码(用虚线框标识)和基于预测树的几何编解码(用点划线框标识)。
对于基于八叉树的几何编码(Octree geometry encoding,OctGeomEnc)而言,基于八叉树的几何编码包括:首先对几何信息进行坐标转换,使点云全都包含在一个包围盒中。然后再进行量化,这一步量化主要起到缩放的作用,由于量化取整,使得一部分点的几何信息相同,根据参数来决定是否移除重复点,量化和移除重复点这一过程又被称为体素化过程。接下来,按照广度优先遍历的顺序不断对包围盒进行树划分(例如八叉树、四叉树、二叉树等),对每个节点的占位码进行编码。在相关技术中,某公司提出了一种隐式几何的划分方式,首先计算点云的包围盒假设dx>dy>dz,该包围盒对应为一个长方体。在几何划分时,首先会基于x轴一直进行二叉树划分,得到两个子节点;直到满足dx=dy>dz条件时,才会基于x和y轴一直进行四叉树划分,得到四个子节点;当最终满足dx=dy=dz条件时,会一直进行八叉树划分,直到划分得到的叶子结点为1×1×1的单位立方体时停止划分,对叶子结点中的点进行编码,生成二进制码流。在基于二叉树/四叉树/八叉树划分的过程中,引入两个参数:K、M。参数K指示在进行八叉树划分之前二叉树/四叉树划分的最多次数;参数M用来指示在进行二叉树/四叉树划分时对应的最小块边长为2M。同时K和M必须满足条件:假设dmax=max(dx,dy,dz),dmin=min(dx,dy,dz),参数K满足:K≥dmax-dmin;参数M满足:M≥dmin。参数K与M之所以满足上述的条件,是因为目前G-PCC在几何隐式划分的过程中,划分方式的优先级为二叉树、四叉树和八叉树,当节点块大小不满足二叉树/四叉树的条件时,才会对节点一直进行八 叉树的划分,直到划分到叶子节点最小单位1×1×1。基于八叉树的几何信息编码模式可以通过利用空间中相邻点之间的相关性来对点云的几何信息进行有效的编码,但是对于一些较为平坦的节点或者具有平面特性的节点,通过利用平面编码可以进一步提升点云几何信息的编码效率。
示例性地,图5A和图5B提供了一种平面位置示意图。其中,图5A示出了一种Z轴方向的低平面位置示意图,图5B示出了一种Z轴方向的高平面位置示意图。如图5A所示,这里的(a)、(a0)、(a1)、(a2)、(a3)均属于Z轴方向的低平面位置,以(a)为例,可以看到当前节点中被占据的四个子节点都位于当前节点在Z轴方向的低平面位置,那么可以认为当前节点属于一个Z平面并且在Z轴方向是一个低平面。同理,如图5B所示,这里的(b)、(b0)、(b1)、(b2)、(b3)均属于Z轴方向的高平面位置,以(b)为例,可以看到当前节点中被占据的四个子节点位于当前节点在Z轴方向的高平面位置,那么可以认为当前节点属于一个Z平面并且在Z轴方向是一个高平面。
进一步地,以图5A中的(a)为例,对八叉树编码和平面编码效率进行比较,图6提供了一种节点编码顺序示意图,即按照图6所示的0、1、2、3、4、5、6、7的顺序进行节点编码。在这里,如果对图5A中的(a)采用八叉树编码方式,那么当前节点的占位信息表示为:11001100。但是如果采用平面编码方式,首先需要编码一个标识符表示当前节点在Z轴方向是一个平面,其次如果当前节点在Z轴方向是一个平面,还需要对当前节点的平面位置进行表示;其次仅仅需要对Z轴方向的低平面节点的占位信息进行编码(即0、2、4、6四个子节点的占位信息),因此基于平面编码方式对当前节点进行编码,仅仅需要编码6个比特(bit),相比相关技术的八叉树编码可以减少2个bit的表示。基于此分析,平面编码相比八叉树编码具有较为明显的编码效率。因此,对于一个被占据的节点,如果在某一个维度上采用平面编码方式进行编码,首先需要对当前节点在该维度上的平面标识(planarMode)和平面位置(PlanePos)信息进行表示,其次基于当前节点的平面信息来对当前节点的占位信息进行编码。示例性地,图7A示出了一种平面标识信息示意图。如图7A所示,这里在Z轴方向为一个低平面;对应地,平面标识信息的取值为真(true)或者1,即planarMode_Z=true;平面位置信息为低平面(low),即PlanePosition_Z=low。图7B示出了另一种平面标识信息示意图。如图7B所示,这里在Z轴方向不为一个平面;对应地,平面标识信息的取值为假(false)或者0,即planarMode_Z=false。
需要注意的是,对于PlaneMode_i:0代表当前节点在i轴方向不是一个平面,1代表当前节点在i轴方向是一个平面。若当前节点在i轴方向是一个平面,则对于PlanePosition_i:0代表当前节点在i轴方向是一个平面,并且平面位置为低平面,1表示当前节点在i轴方向上是一个高平面。其中,i表示坐标维度,可以为X轴方向、Y轴方向或者Z轴方向,故i=0,1,2。
在G-PCC标准中,判断一个节点是否满足平面编码的条件以及在该节点满足平面编码条件时,需要对该节点的平面标识和平面位置信息的预测编码。
在本申请实施例中,当前G-PCC标准中存在三种判断节点是否满足平面编码的判断条件,下面对其逐一进行详细说明。
一、根据节点在每个维度上的平面概率进行判断。
(1)确定当前节点的局部区域密度(local_node_density);
(2)确定当前节点在每个维度上的概率Prob(i)。
在节点的局部区域密度小于阈值Th(例如Th=3)时,利用当前节点在三个坐标维度上的平面概率Prob(i)和阈值Th0、Th1和Th2进行比较,其中Th0<Th1<Th2(例如,Th0=0.6,Th1=0.77,Th2=0.88),这里可以利用Eligiblei(i=0,1,2)表示每个维度上是否启动平面编码:Eligiblei=Prob(i)>=threshold。
需要注意的是,threshold是进行自适应变化的,例如,当Prob(0)>Prob(1)>Prob(2)时,则Eligiblei的设置如下:
Eligible0=Prob(0)>=Th0;
Eligible1=Prob(1)>=Th1;
Eligible2=Prob(2)>=Th2。
当Prob(1)>Prob(0)>Prob(2)时,则Eligiblei的设置如下:
Eligible0=Prob(0)>=Th1;
Eligible1=Prob(1)>=Th0;
Eligible2=Prob(2)>=Th2。
在这里,Prob(i)的更新具体如下:
Prob(i)new=(L×Prob(i)+δ(coded node))/L+1        (1)
其中,L=255;另外,若coded node节点是一个平面,则δ(coded node)为1;否则δ(coded node)为0。
在这里,local_node_density的更新具体如下:
local_node_density new=local_node_density+4*numSiblings       (2)
其中,local_node_density初始化为4,numSiblings为该节点的兄弟姐妹节点数目。示例性地,图8示出了一种当前节点的兄弟姐妹节点示意图。如图8所示,当前节点为用斜线填充的节点,用网格填充的节点为兄弟姐妹节点,那么当前节点的兄弟姐妹节点数目为5(包括当前节点自身)。
二、根据当前层的点云密度来判断当前层节点是否满足平面编码。
利用当前层点的密度来判断是否对当前层的节点进行平面编码。假设当前待编码点云的点数为pointCount,经过推断直接编码模式(Infer Direct Coding Model,IDCM)编码已经重建出的点数为numPointCountRecon,又因为八叉树是基于广度优先遍历的顺序进行编码,因此可以得到当前层待编码的节点数目假设为nodeCount,那么判断当前层是否启动平面编码假设为planarEligibleKOctreeDepth,具体为:planarEligibleK OctreeDepth=(pointCount-numPointCountRecon)<nodeCount×1.3。
其中,若(pointCount-numPointCountRecon)小于nodeCount×1.3,则planarEligibleK OctreeDepth为true;若(pointCount-numPointCountRecon)不小于nodeCount×1.3,则planarEligibleKOctreeDepth为false。这样,当planarEligibleKOctreeDepth为true时,则在当前层所有节点都进行平面编码;否则在当前层所有节点都不进行平面编码,仅仅采用八叉树编码。
三、根据激光雷达点云的采集参数来判断当前节点是否满足平面编码。
图9示出了一种激光雷达与节点的相交示意图。如图9所示,用网格填充的节点同时被两个激光射线(Laser)穿过,因此当前节点在Z轴垂直方向上不是一个平面;用斜线填充的节点足够小到不能同时被两个Laser同时穿过,因此斜线填充的节点在Z轴垂直方向上有可能是一个平面。
进一步地,针对满足平面编码条件的节点,可以对平面标识信息和平面位置信息进行预测编码。
首先,平面标识信息的预测编码。
在这里,仅仅采用三个上下文信息进行编码,即各个坐标维度上的平面标识分开进行上下文设计。
其次,平面位置信息的预测编码。
应理解,针对非激光雷达点云平面位置信息的编码而言,平面位置信息的预测编码可以包括:
(a)利用邻域节点的占位信息进行预测得到当前节点的平面位置信息为三元素:预测为低平面、预测为高平面和无法预测;
(b)与当前节点在相同划分深度以及相同坐标下的节点与当前节点之间的空间距离:“近”和“远”;
(c)与当前节点在相同划分深度以及相同坐标下的节点如果是一个平面,则确定该节点的平面位置;
(d)坐标维度(i=0,1,2)。
需要说明的是,在本申请实施例中,确定出与当前节点在相同划分深度以及相同坐标下的节点和当前节点之间的空间距离之后,如果该空间距离小于预设距离阈值,那么可以确定该空间距离为“近”;或者,如果该空间距离大于预设距离阈值,那么可以确定该空间距离为“远”。
示例性地,图10示出了一种处于相同划分深度以及相同坐标的邻域节点示意图。如图10所示,加粗的大立方体表示父节点(Parent node),其内部网格填充的小立方体表示当前节点(Current node),并且示出了当前节点的交点位置(Vertex position);白色填充的小立方体表示处于相同划分深度以及相同坐标的邻域节点,当前节点与邻域节点之间的距离为空间距离,可以判断为“近”或“远”;另外,如果该邻域节点为一个平面,那么还需要该邻域节点的平面位置(Planar position)。
这样,如图10所示,当前节点为网格填充的小立方体,则在相同的八叉树划分深度等级下,以及相同的垂直坐标下查找邻域节点为白色填充的小立方体,判断两个节点之间的距离为“近”和“远”,并且参考节点的平面位置。
进一步地,在本申请实施例中,图11示出了一种当前节点位于父节点的低平面位置示意图。如图11所示,(a)、(b)、(c)示出了三种当前节点位于父节点的低平面位置的示例。具体说明如下:
①如果点填充节点的子节点4到7中有任何一个被占用,而所有网格填充节点都未被占用,则极有可能在当前节点(用斜线填充)中存在一个平面,且该平面位置较低。
②如果点填充节点的子节点4到7都未被占用,而任何网格填充节点被占用,则极有可能在当前节点(用斜线填充)中存在一个平面,且该平面位置较高。
③如果点填充节点的子节点4到7均为空节点,网格填充节点均为空节点,则无法推断平面位置,故标记为未知。
④如果点填充节点的子节点4到7中有任何一个被占用,而网格填充节点中有任何一个被占用,此时也无法推断出平面位置,因此将其标记为未知。
在本申请实施例中,图12示出了一种当前节点位于父节点的高平面位置示意图。如图12所示,(a)、(b)、(c)示出了三种当前节点位于父节点的高平面位置的示例。具体说明如下:
①如果网格填充节点的子节点4到7中有任何一个节点被占用,而点填充节点未被占用,则极有可能在当前节点(用斜线填充)中存在一个平面,且平面位置较低。
②如果网格填充节点的子节点4到7均未被占用,而点填充节点被占用,则极有可能在当前节点(用斜线填充)中存在平面,且平面位置较高。
③如果网格填充节点的子节点4到7都是未被占用的,而点填充节点是未被占用的,此时无法推断平面位置,因此标记为未知。
④如果网格填充节点的子节点4到7中有一个被占用,而点填充节点被占用,此时无法推断平面位置,因此标记为未知。
还应理解,针对激光雷达点云平面位置信息的编码而言,图13示出了一种激光雷达点云平面位置信息的预测编码示意图。如图13所示,在激光雷达的发射角度为θbottom时,这时候可以映射为低平面(Bottom virtual plane);在激光雷达的发射角度为θtop时,这时候可以映射为高平面(Top virtual plane)。
也就是说,通过利用激光雷达采集参数来预测当前节点的平面位置,通过利用当前节点与激光射线相交的位置来将位置量化为多个区间,最终作为当前节点平面位置的上下文信息。具体计算过程如下:假设激光雷达的坐标为(xLidar,yLidar,zLidar),当前节点的几何坐标为(x,y,z),那么首先计算当前节点相对于激光雷达的垂直正切值tanθ,计算公式如下:
进一步地,又因为每个Laser会相对于激光雷达有一定偏移角度,因此还需要计算当前节点相对于Laser的相对正切值tanθcorr,L,具体计算如下:
最终会利用当前节点的相对正切值tanθcorr,L来对当前节点的平面位置进行预测,具体如下,假设当前节点下边界的正切值为tan(θbottom),上边界的正切值为tan(θtop),根据tanθcorr,L将平面位置量化为4个量化区间,即确定平面位置的上下文信息。
但是基于八叉树的几何信息编码模式仅对空间中具有相关性的点有高效的压缩速率,而对于在几何空间中处于孤立位置的点来说,使用直接编码模式(Direct Coding Model,DCM)可以大大降低复杂度。对于八叉树中的所有节点,DCM的使用不是通过标志位信息来表示的,而是通过当前节点父节点和邻居信息来进行推断得到。判断当前节点是否具有DCM编码资格的方式有三种,具体如下:
(1)当前节点没有兄弟姐妹子节点,即当前节点的父节点只有一个孩子节点,同时当前节点父节点的父节点仅有两个被占据子节点,即当前节点最多只有一个邻居节点。
(2)当前节点的父节点仅有当前节点一个占据子节点,同时与当前节点共用一个面的六个邻居节点也都属于空节点。
(3)当前节点的兄弟姐妹节点数目大于1。
示例性地,图14提供了一种IDCM编码示意图。如果当前节点不具有DCM编码资格将对其进行八叉树划分,若具有DCM编码资格将进一步判断该节点中包含的点数,当点数小于阈值(例如2)时,则对该节点进行DCM编码,否则将继续进行八叉树划分。当应用DCM编码模式时,首先需要编码当前节点是否是一个真正的孤立点,即IDCM_flag,当IDCM_flag为true时,则当前节点采用DCM编码,否则仍然采用八叉树编码。在当前节点满足DCM编码时,需要编码当前节点的DCM编码模式,目前存在两种DCM模式,分别是:(a)仅仅只有一个点存在(或者是多个点,但是属于重复点);(b)含有两个点。最后需要编码每个点的几何信息,假设节点的边长为2d时,对该节点几何坐标的每一个分量进行编码时需要d比特,该比特信息直接被编进码流中。这里需要注意的是,在对激光雷达点云进行编码时,通过利用激光雷达采集参数来对三个维度的坐标信息进行预测编码,从而可以进一步提升几何信息的编码效率。
进一步地,下面针对IDCM编码的过程进行详细介绍。
当前节点满足DCM编码模式时,首先编码当前节点的点数目numPoints;根据不同的DirectMode来对当前节点的点数目进行编码:
(1)如果当前节点不满足DCM节点的要求,则直接退出(即点数大于2个点,并且不是重复点)。
(2)当前节点含有的点数numPonts小于或等于2,则编码过程如下:
i)首先编码当前节点的numPonts是否大于1;
ii)如果当前节点只有一个点并且几何编码环境为几何无损编码,则需要编码当前节点的第二个点不是重复点。
(3)当前节点含有的点数numPonts大于2,则编码过程如下:
i)首先编码当前节点的numPonts小于或等于1;
ii)其次编码当前节点的第二个点是一个重复点,其次编码当前节点的重复点数目是否大于1,当重复点数目大于1时,需要对剩余的重复点数目进行指数哥伦布解码。
在编码完成当前节点的点数目之后,对当前节点中包含点的坐标信息进行编码。下面将分别对激光雷达点云和面向人眼点云进行详细介绍。
(一)面向人眼点云。
(1)如果当前节点中仅仅只含有一个点,则会对点的三个维度方向的几何信息进行直接编码(Bypass coding);
(2)如果当前节点中含有两个点,则会首先通过利用点的几何坐标得到优先编码的坐标轴dirextAxis。这里需要注意的是,目前比较的坐标轴只包含x轴和y轴,不包含z轴。假设当前节点的几何坐标为nodePos,则判断的方式如下:
dirextAxis=!(nodePos[0]<nodePos[1])      (5)
也就是会将节点坐标几何位置小的轴作为优先编码的坐标轴dirextAxis,其次按照如下方式首先对优先编码的坐标轴dirextAxis几何信息进行编码。假设优先编码的轴对应的代编码几何bit深度为nodeSizeLog2,并假设两个点的坐标分别为pointPos[0]和pointPos[1]。具体编码过程如下:
在编码完成优先编码的坐标轴dirextAxis之后,再继续对当前节点的几何坐标进行直接编码。假设每个点的剩余编码bit深度为nodeSizeLog2,则具体编码过程如下:
for(int axisIdx=0;axisIdx<3;++axisIdx)
for(int mask=(1<<nodeSizeLog2[axisIdx])>>1;mask;mask>>1)
             encodePosBit(!!(pointPos[axisIdx]&mask))。
(二)面向激光雷达点云。
如果当前节点中含有两个点,则会首先通过利用点的几何坐标得到优先编码的坐标轴dirextAxis,假设当前节点的几何坐标为nodePos,则判断的方式如下:
dirextAxis=!(nodePos[0]<nodePos[1])
也就是会将节点坐标几何位置小的轴作为优先编码的坐标轴dirextAxis,这里需要注意的是,目前比较的坐标轴只包含x轴和y轴,不包含z轴。其次按照如下方式首先对优先编码的坐标轴dirextAxis几何信息进行编码,假设优先编码的轴对应的代编码几何bit深度为nodeSizeLog2,并假设两个点的坐标分别为pointPos[0]和pointPos[1]。具体编码过程如下:
在编码完成优先编码的坐标轴dirextAxis之后,再对当前节点的几何坐标进行编码。
由于激光雷达点云可以得到激光雷达点云的采集参数,通过利用可以预测当前节点的几何坐标信息,从而可以进一步提升点云的几何信息编码效率。同样的首先利用当前节点的几何信息nodePos得到一个直接编码的主轴方向,其次利用已经完成编码的方向的几何信息来对另外一个维度的几何信息进行预测编码。同样假设直接编码的轴方向是directAxis,并且假设直接编码中的代编码bit深度为nodeSizeLog2, 则编码方式如下:
for(int mask=(1<<nodeSizeLog2)>>1;mask;mask>>1);
            encodePosBit(!!(pointPos[directAxis]&mask))。
这里需要注意的是,在这里会将directAxis方向的几何精度信息全部编码。
示例性地,图15提供了一种旋转激光雷达获取点云的坐标转换示意图。其中,在笛卡尔坐标系下,对于每一个节点的(x,y,z)坐标,均可以转换为用(R,i)表示。另外,激光扫描器(Laser Scanner)可以按照预设角度进行激光扫描,在i的不同取值下,可以得到不同的θ(i)。例如,在i等于1时,这时候可以得到θ(1),对应的扫描角度为-15°;在i等于2时,这时候可以得到θ(2),对应的扫描角度为-13°;在i等于10时,这时候可以得到θ(10),对应的扫描角度为+13°;在i等于9时,这时候可以得到θ(19),对应的扫描角度为+15°。
这样,在编码完成directAxis坐标方向的所有精度之后,会首先计算当前点所对应的LaserIdx,即图15中的pointLaserIdx号,并且计算当前节点的LaserIdx,即nodeLaserIdx;其次会利用节点的LaserIdx即nodeLaserIdx来对点的LaserIdx即pointLaserIdx进行预测编码,其中节点或者点的LaserIdx的计算方式如下。假设点的几何坐标为pointPos,激光射线的起始坐标为LidarOrigin,并且假设Laser的数目为LaserNum,每个Laser的正切值为tanθi,每个Laser在垂直方向上的偏移位置为Zi,则:
在计算得到当前点的LaserIdx之后,首先会利用当前节点的LaserIdx对点的pointLaserIdx进行预测编码。在编码完成当前点的LaserIdx之后,对当前点三个维度的几何信息利用激光雷达的采集参数进行预测编码。
示例性地,图16示出了一种X轴或Y轴方向的预测编码示意图。如图16所示,用网格填充的方框表示当前点(Current node),用斜线填充的方框表示已编码点(Already coded node)。在这里,首先利用当前点对应的LaserIdx得到对应的水平方位角的预测值,即其次利用当前点对应的节点几何信息得到节点对应的水平方位角度其中,假设节点的几何坐标为nodePos,则水平方位角与节点几何信息之间的计算方式如下:
通过利用激光雷达的采集参数,可以得到每个Laser的旋转点数numPoints,即代表每个激光射线旋转一圈得到的点数,则可以利用每个Laser的旋转点数计算得到每个Laser的旋转角速度deltaPhi,计算方式如下:
进一步地,利用节点的水平方位角以及当前点对应的Laser前一个编码点的水平方位角计算得到当前点对应的水平方位角预测值即如图17A和图17B所示的水平方位角的预测值。其中,图17A示出了一种通过水平方位角来进行预测Y平面的角度示意图,图17B示出了一种通过水平方位角来进行预测X平面的角度示意图。在这里,对于当前点对应的水平方位角预测值计算方式如下:
示例性地,图18示出了另一种X轴或Y轴方向的预测编码示意图。如图18所示,用网格填充的部分(左侧)表示低平面,用点填充的部分(右侧)表示高平面,表示当前节点的低平面水平方位角,表示当前节点的高平面水平方位角,表示当前节点对应的水平方位角预测值。
这样,通过利用水平方位角的预测值以及当前节点的低平面水平方位角和高平面水平方位角来对当前节点的几何信息进行预测编码。具体如下所示:


int context=(angLel≥0&&angLeR≥0)||(angLel<0&&angLeR<0)?0:2;
int minAngle=std∷min(abs(angLel),abs(angLeR));
int maxAngle=std∷max(abs(angLel),abs(angLeR));
context+=maxAngle>minAngle?0:1;
context+=maxAngle>minAngle?0:4。
在编码完成点的LaserIdx之后,会利用当前点所对应的LaserIdx对当前点的Z轴方向进行预测编码,即当前通过利用当前点的x和y信息计算得到雷达坐标系的深度信息radius,其次利用当前点的激光LaserIdx得到当前点的正切值以及垂直方向的偏移量,则可以得到当前点的Z轴方向的预测值即Z_pred。具体如下所示:

int tanTheta=tanθlaserIdx
int zOffset=ZlaserIdx
Z_pred=radius×tanTheta-zOffset。
进一步地,利用Z_pred对当前点的Z轴方向的几何信息进行预测编码得到预测残差Z_res,最终对Z_res进行编码。
需要注意的是,在节点划分到叶子节点时,在几何无损编码的情况下,需要对叶子节点中的重复点数目进行编码。最终对所有节点的占位信息进行编码,生成二进制码流。另外G-PCC目前引入了一种平面编码模式,在对几何进行划分的过程中,会判断当前节点的子节点是否处于同一平面,如果当前节点的子节点满足同一平面的条件,会用该平面对当前节点的子节点进行表示。
对于基于八叉树的几何解码而言,解码端按照广度优先遍历的顺序,在对每个节点的占位信息解码之前,首先会利用已经重建得到的几何信息来判断当前节点是否进行平面解码或者IDCM解码,如果当前节点满足平面解码的条件,则会首先对当前节点的平面标识和平面位置信息进行解码,其次基于平面信息来对当前节点的占位信息进行解码;如果当前节点满足IDCM解码的条件,则会首先解码当前节点是否是一个真正的IDCM节点,如果是一个真正的IDCM解码,则会继续解析当前节点的DCM解码模式,其次可以得到当前DCM节点中的点数目,最后对每个点的几何信息进行解码。对于既不满足平面解码也不满足DCM解码的节点,会对当前节点的占位信息进行解码。通过按照这样的方式不断解析得到每个节点的占位码,并且依次不断划分节点,直至划分得到1×1×1的单位立方体时停止划分,解析得到每个叶子节点中包含的点数,最终恢复得到几何重构点云信息。
下面对IDCM解码的过程进行详细介绍。
与编码端的处理过程类似,首先利用先验信息来决定节点是否启动IDCM,即IDCM的启动条件如下:
(1)当前节点没有兄弟姐妹子节点,即当前节点的父节点只有一个孩子节点,同时当前节点父节点的父节点仅有两个被占据子节点,即当前节点最多只有一个邻居节点。
(2)当前节点的父节点仅有当前节点一个占据子节点,同时与当前节点共用一个面的六个邻居节点也都属于空节点。
(3)当前节点的兄弟姐妹节点数目大于1。
进一步地,当节点满足DCM编码的条件时,首先解码当前节点是否是一个真正的DCM节点,即IDCM_flag;当IDCM_flag为true时,则当前节点采用DCM编码,否则仍然采用八叉树编码。
其次解码当前节点的点数目numPoints,具体的解码方式如下所示:
i)首先解码当前节点的numPonts是否大于1;
ii)如果解码得到当前节点的numPonts大于1,则继续解码第二个点是否是一个重复点;如果第二个点不是重复点,则这里可以隐性推断出满足DCM模式的第二种,只含有两个点;
iii)如果解码得到当前节点的numPonts小于等于1,则继续解码第二个点是否是一个重复点;如果第二个点不是重复点,则这里可以隐性推断出满足DCM模式的第二种,只含有一个点;如果解码得到第二个点是一个重复点,则可以推断出满足DCM模式的第三种,含有多个点,但是都是重复点,则继续解码重复点的数目是否大于1(熵解码),如果大于1,则继续解码剩余重复点的数目(利用指数哥伦布进行解码)。
如果当前节点不满足DCM节点的要求,则直接退出(即点数大于2个点,并且不是重复点)。
在解码完成当前节点的点数目之后,对当前节点中包含点的坐标信息进行解码。下面将分别对激光雷达点云和面向人眼点云进行详细介绍。
(一)面向人眼点云。
(1)如果当前节点中仅仅只含有一个点,则会对点的三个维度方向的几何信息进行直接解码(Bypass coding);
(2)如果当前节点中含有两个点,则会首先通过利用点的几何坐标得到优先解码的坐标轴dirextAxis,这里需要注意的是,目前比较的坐标轴只包含x和y轴,不包含z轴。假设当前节点的几何坐标为nodePos,则判断的方式如下:
dirextAxis=!(nodePos[0]<nodePos[1])       (9)
也就是会将节点坐标几何位置小的轴作为优先解码的坐标轴dirextAxis,其次按照如下方式首先对优先解码的坐标轴dirextAxis几何信息进行解码。假设优先解码的轴对应的待解码几何bit深度为nodeSizeLog2,并假设两个点的坐标分别为pointPos[0]和pointPos[1]。具体编码过程如下:
在解码完成优先解码的坐标轴dirextAxis之后,再继续对当前点的几何坐标进行直接解码。假设每个点的剩余编码bit深度为nodeSizeLog2,并假设点的坐标信息为pointPos,则具体解码过程如下:
(二)面向激光雷达点云。
如果当前节点中含有两个点,则会首先通过利用点的几何坐标得到优先解码的坐标轴dirextAxis,假设当前节点的几何坐标为nodePos,则判断的方式如下:
dirextAxis=!(nodePos[0]<nodePos[1])         (10)
也就是会将节点坐标几何位置小的轴作为优先解码的坐标轴dirextAxis,这里需要注意的是,目前比较的坐标轴只包含x轴和y轴,不包含z轴。其次按照如下方式首先对优先编码的坐标轴dirextAxis几何信息进行解码,假设优先解码的轴对应的代编码几何bit深度为nodeSizeLog2,并假设两个点的坐标分别为pointPos[0]和pointPos[1]。具体编码过程如下:

在解码完优先解码的坐标轴dirextAxis之后,再对当前点的几何坐标进行解码。
同样的首先利用当前节点的几何信息nodePos得到一个直接解码的主轴方向,其次利用已经完成解码的方向的几何信息来对另外一个维度的几何信息进行解码。同样假设直接解码的轴方向是directAxis,并且假设直接解码中的待解码bit深度为nodeSizeLog2,则解码方式如下:
这里需要注意的是,在这里会将directAxis方向的几何精度信息全部解码。
在解码完成directAxis坐标方向的所有精度之后,会首先计算当前节点的LaserIdx,即nodeLaserIdx;其次会利用节点的LaserIdx即nodeLaserIdx来对点的LaserIdx即pointLaserIdx进行预测解码,其中节点或者点的LaserIdx的计算方式跟编码端相同。最终对当前点的LaserIdx与节点的LaserIdx预测残差信息进行解码得到ResLaserIdx,则解码方式如下:
PointLaserIdx=nodeLaserIdx+ResLaserIdx        (11)
在解码完成当前点的LaserIdx之后,对当前点三个维度的几何信息利用激光雷达的采集参数进行预测解码。具体算法如下:
如图11所示,首先利用当前点对应的LaserIdx得到对应的水平方位角的预测值,即其次利用当前点对应的节点几何信息得到节点对应的水平方位角度其中,假设节点的几何坐标为nodePos,则水平方位角与节点几何信息之间的计算方式如下:
通过利用激光雷达的采集参数,可以得到每个Laser的旋转点数numPoints,即代表每个激光射线旋转一圈得到的点数,则可以利用每个Laser的旋转点数计算得到每个Laser的旋转角速度deltaPhi,计算方式如下:
进一步地,利用节点的水平方位角以及当前点对应的Laser前一个编码点的水平方位角计算得到当前点对应的水平方位角预测值即如图17A和图17B所示的水平方位角的预测值。计算方式如下:
这样,通过利用水平方位角的预测值以及当前节点的低平面水平方位角和高平面的水平方位角来对当前节点的几何信息进行预测解码。具体如下所示:


int context=(angLel≥0&&angLeR≥0)||(angLel<0&&angLeR<0)?0:2;
int absAngleL=abs(angLel);
int absAngleR=abs(angLeR);
context+=absAngleL>absAngleR?0:1;
context+=maxAngle>minAngle<<1?4:0。
在解码完成点的LaserIdx之后,会利用当前点所对应的LaserIdx对当前点的Z轴方向进行预测解码,即当前通过利用当前点的x和y信息计算得到雷达坐标系的深度信息radius,其次利用当前点的激光LaserIdx得到当前点的正切值以及垂直方向的偏移量,则可以得到当前点的Z轴方向的预测值即Z_pred。具体如下所示:

int tanTheta=tanθlaserIdx
int zOffset=ZlaserIdx
Z_pred=radius×tanTheta-zOffset。
进一步地,利用解码得到的Z_res和Z_pred来重建恢复得到当前点Z轴方向的几何信息。
对于基于三角面片集(triangle soup,trisoup)的几何信息编码而言,在基于trisoup的几何信息编码框架中,同样也要先进行几何划分,但区别于基于二叉树/四叉树/八叉树的几何信息编码,该方法不 需要将点云逐级划分到边长为1×1×1的单位立方体,而是划分到子块(block)边长为W时停止划分,基于每个block中点云的分布所形成的表面,得到该表面与block的十二条边所产生的至多十二个交点(vertex)。依次编码每个block的vertex坐标,生成二进制码流。
对于基于trisoup的点云几何信息重建而言,在解码端进行点云几何信息重建时,首先解码vertex坐标用于完成三角面片重建,该过程如图19A、图19B和图19C所示。其中,图19A所示的block中存在3个交点(v1,v2,v3),利用这3个交点按照一定顺序所构成的三角面片集被称为triangle soup,即trisoup,如图19B所示。之后,在该三角面片集上进行采样,将得到的采样点作为该block内的重建点云,如图19C所示。
对于基于预测树的几何编码(Predictive geometry coding,PredGeomTree)而言,基于预测树的几何编码包括:首先对输入点云进行排序,目前采用的排序方法包括无序、莫顿序、方位角序和径向距离序。在编码端通过利用两种不同的方式建立预测树结构,其中包括:KD-Tree(高时延慢速模式)和低时延快速模式(利用激光雷达标定信息)。在利用激光雷达标定信息时,将每个点划分到不同的Laser上,按照不同的Laser建立预测树结构。接下来基于预测树的结构,遍历预测树中的每个节点,通过选取不同的预测模式对节点的几何位置信息进行预测得到预测残差,并且利用量化参数对几何预测残差进行量化。最终通过不断迭代,对预测树节点位置信息的预测残差、预测树结构以及量化参数等进行编码,生成二进制码流。
对于基于预测树的几何解码而言,解码端通过不断解析码流,重构预测树结构,其次通过解析得到每个预测节点的几何位置预测残差信息以及量化参数,并且对预测残差进行反量化,恢复得到每个节点的重构几何位置信息,最终完成解码端的几何重构。
在几何编码完成后,需要对几何信息进行重建。目前,属性编码主要针对颜色信息进行。首先,将颜色信息从RGB颜色空间转换到YUV颜色空间。然后,利用重建的几何信息对点云重新着色,使得未编码的属性信息与重建的几何信息对应起来。在颜色信息编码中,主要有两种变换方法,一是依赖于LOD划分的基于距离的提升变换,二是直接进行RAHT变换,这两种方法都会将颜色信息从空间域转换到频域,通过变换得到高频系数和低频系数,最后对系数进行量化并编码,生成二进制码流,具体参见图4A和图4B所示。
进一步地,在利用几何信息来对属性信息进行预测时,可以利用莫顿码进行最近邻居搜索,点云中每点对应的莫顿码可以由该点的几何坐标得到。计算莫顿码的具体方法描述如下所示,对于每一个分量用d比特二进制数表示的三维坐标,其三个分量可以表示为:
其中,分别是x,y,z的最高位到最低位对应的二进制数值。莫顿码M是对x,y,z从最高位开始,依次交叉排列到最低位,M的计算公式如下所示:
其中,分别是M的最高位到最低位的值。在得到点云中每个点的莫顿码M后,将点云中的点按莫顿码由小到大的顺序进行排列,并将每个点的权重值w设为1。
还可以理解,对于G-PCC编解码框架而言,通用测试条件如下:
(1)测试条件共4种:
条件1:几何位置有限度有损、属性有损;
条件2:几何位置无损、属性有损;
条件3:几何位置无损、属性有限度有损;
条件4:几何位置无损、属性无损。
(2)通用测试序列包括Cat1A,Cat1B,Cat3-fused,Cat3-frame共四类,其中Cat2-frame点云只包含反射率属性信息,Cat1A、Cat1B点云只包含颜色属性信息,Cat3-fused点云同时包含颜色和反射率属性信息。
(3)技术路线:共2种,以几何压缩所采用的算法进行区分。
技术路线1:八叉树编码分支。
在编码端,将包围盒依次划分得到子立方体,对非空的(包含点云中的点)的子立方体继续进行划分,直到划分得到的叶子结点为1×1×1的单位立方体时停止划分,在几何无损编码情况下,需要对叶子节点中所包含的点数进行编码,最终完成几何八叉树的编码,生成二进制码流。
在解码端,解码端按照广度优先遍历的顺序,通过不断解析得到每个节点的占位码,并且依次不断划分节点,直至划分得到1×1×1的单位立方体时停止划分,在几何无损解码的情况下,需要解析得到每个叶子节点中包含的点数,最终恢复得到几何重构点云信息。
技术路线2:预测树编码分支。
在编码端,通过利用两种不同的方式建立预测树结构,其中包括:基于KD-Tree(高时延慢速模式)和利用激光雷达标定信息(低时延快速模式),利用激光雷达标定信息,可以将每个点划分到不同的Laser上,按照不同的Laser建立预测树结构。接下来基于预测树的结构,遍历预测树中的每个节点,通过选取不同的预测模式对节点的几何位置信息进行预测得到预测残差,并且利用量化参数对几何预测残差进行量化。最终通过不断迭代,对预测树节点位置信息的预测残差、预测树结构以及量化参数等进行编码,生成二进制码流。
在解码端,解码端通过不断解析码流,重构预测树结构,其次通过解析得到每个预测节点的几何位置预测残差信息以及量化参数,并且对预测残差进行反量化,恢复得到每个节点的重构几何位置信息,最终完成解码端的几何重构。
还需要说明的是,在如图4A或图4B所示,目前G-PCC编码框架包含三种属性编码方法:预测变换(Predicting Transform,PT)、提升变换(Lifting Transform,LT)以及区域自适应分层变换(Region Adaptive Hierarchical Transform,RAHT)。其中,前两者是以LOD的生成顺序为依据对点云预测编码,RAHT则是依据八叉树的构建层级自下而上对属性信息进行自适应变换。下面将分别对这三种点云属性编码方法进行具体介绍。
(a)点云属性信息的预测编码。
目前G-PCC的属性预测模块采用一种基于分层(Level-of-details,LoDs)结构的最近邻属性预测编码方案,LOD的构造方法包括基于距离的LOD构造方案、基于固定采样率的LOD构造方案以及基于八叉树的LOD构造方案等。在基于距离阈值的LOD构造方案中,构造LOD之前首先对点云进行Morton排序,来保证相邻点之间具有较强的属性相关性。图20为一种基于距离的LOD构造过程的示意图。如图20所示,根据用户提前预设的L个曼哈顿(Manhattan)距离(dl),l=0,1,…L-1;将点云划分成L个不同的点云细节层(Rl),l=0,1,…L-1,其中(dl)l=0,1,…L-1满足dl<dl-1。LOD的构造过程如下所述:
(1)首先将点云中所有点都标记为未访问过,建立一个集合V用来存储已经访问过的点集;(2)对于每一次迭代l,通过对点云中的点进行遍历,如果当前点已经被访问过,则忽略该点,否则计算当前点到点集V的最小距离D,如果D<dl,则忽略该点;否则将当前点标记为已访问并将当前点加入细化层Rl和点集V;(3)细节层次LODl中的点由细化层R0,R1,R2…Rl中的点构成;(4)不断重复上述步骤,直至所有的点都被标记为已访问。
在LOD的结构基础上,每个点的属性值通过利用同一层或更高一层LOD中点的属性重建值进行线性加权预测,其中参考预测邻居的最大数目由编码器高层语法元素决定。对于每个点的属性,在编码端利用率失真优化算法选取通过利用搜索到的N个最近邻点的属性进行加权预测或者选择单个最近邻点的属性进行预测,最后对选取的预测模式以及预测残差进行编码。
其中,N代表点i最近邻点集中预测点的数目,Pi代表点i的N个最近邻点的合,Dm代表了最近邻点m到当前点i的空间几何距离,Attrm代表了最近邻点m重建之后的属性值,Attri′代表了对当前点i的属性预测值,点数N为提前预设的数值。
为了权衡属性编码效率和不同LOD层之间的并行处理,在编码器高层语法元素引入了一个开关可以控制是否引入LOD层内预测。如果开启则启动LOD层内预测,可以利用同一LOD层内的点进行预测。需要注意的是,当LOD层的数目为1时,总是使用LOD层内预测。
图21为一种LOD生成过程的可视化结果示意图。如图21所示,这里提供了一种基于距离的LOD生成过程的主观示例。具体是(从左向右):第一层中的点是代表点云的外轮廓;随着细节层的增加,点云细节描述逐渐清晰。
图22为一种属性预测的编码流程示意图。如图22所示,针对G-PCC属性预测的具体流程,对于原始点云,首先搜索第K个点的三个近邻点,然后进行属性预测;根据第K个点的属性预测值与第K个点的属性原始值进行差值计算,可以得到第K个点的预测残差;然后进行量化与算术编码,最终生成属性码率。
(i)最优预测值选取:
LOD构建完成以后,根据LOD的生成顺序,首先从已编码的数据点中找到当前待编码点的三个最近邻点。将这3个最近邻点的属性重建值,作为当前待编码点的候选预测值;然后,根据率失真优化(Rate-Distortion Optimal,RDO)从中选择最优的预测值。例如,当编码图20中点P2的属性值时,将最近邻居点P4属性值的预测变量索引设为1;将次近邻点P5和三近邻点P0的属性预测变量索引分别设为2和3;将点P0、P5和P4的加权平均值的预测变量索引设为0,如表1所示;最后,利用RDO 选择最佳预测变量。其中加权平均的公式如下所示:
其中,表示近邻点j到当前点i的空间几何权重:
表示对当前点i的属性预测值,j表示3个近邻点的索引,代表了近邻点重建之后的属性值,xi,yi,zi是当前点i的几何位置坐标,xij,yij,zij为近邻点j的几何坐标。
示例性地,表1提供了一种属性编码的候选预测项样本示例。
表1
(ii)属性预测残差及量化:
通过上述预测得到当前点i的属性预测值(k为点云的总点数)。令(ai)i∈0…k-1为当前点的属性原始值,则属性残差(ri)i∈0…k-1记为:
进一步对预测残差进行量化:
其中,Qi表示当前点i的量化后的属性残差,Qs为量化步长(Quantization step,Qs),可以由CTC规定的量化参数QP(Quantization Parameter,QP)计算得出。
(iii)编码端重建属性值:
编码端重建的目的是为了后续点的预测。在重建属性值之前要对残差进行反量化,记为反量化后的残差:
与预测值相加得到点i的重建值
在基于LOD划分的基础上进行属性最近邻查找时,目前存在两大类算法:帧内最近邻查找和帧间最近邻查找。其中,帧间的最近邻查找算法具体如下,帧内的最近邻查找可以分为层间最近邻查找和层内最近邻查找两种算法。
(i)帧内最近邻查找:
帧内最近邻查找分为层间最近邻查找和层内最近邻查找两种算法。LOD划分之后,类似一个金字塔结构,如图23所示。
在一种具体的实现方式中,对于层间最近邻查找,金字塔结构如图24所示。图25为一种层间最
近邻查找的LOD构造过程示意图。如图25所示,基于几何信息划分得到不同的LOD层,得到
LOD0、LOD1和LOD2,利用LOD0中的点去预测下一层LOD中点的属性在层间最近邻查找的
过程中。
下面将对帧内最近邻查找的整个过程进行详细地介绍。
在整个LOD的划分过程中,存在三个集合O(k)、L(k)以及I(k)。其中,k为LOD划分时LOD层的索引,I(k)为当前LOD层划分时的输入点集,经过LOD划分,得到O(k)集合以及L(k)集合,O(k)集合存储的是采样点集,L(k)为当前LOD层中的点集。即整个LOD划分的过程如下:
(1)初始化。
if k=0,L(k)←{};否则,L(k)←L(k-1);
O(k)←{};
(2)利用LOD划分算法,将采样点存入O(k),其余的点划分到L(k);
(3)进行下一次迭代时L←O(k)。
这里需要注意的是,由于整个LOD划分的过程是基于莫顿码进行划分的,因此O(k)、L(k)以及I(k)存储的是点对应的莫顿码索引。
在进行层间最近邻查找时,即L(k)集合中的点在O(k)集合中进行最近邻查找,具体的查找算法如 下:
以基于空间关系进行最近邻查找为例,在对当前点P进行预测时,通过利用点P对应的父块(Block B)进行邻居搜索,如图26所示,搜索与当前父块共面、共线邻居块内的点来进行属性预测。
其中,图27A示出了一种共面的空间关系示意图,这里共有6个与当前父块具有关系的空间块。图27B示出了一种共面和共线的空间关系示意图,这里共有18个与当前父块具有关系的空间块。图27C示出了一种共面、共线和共点的空间关系示意图,这里共有26个与当前父块具有关系的空间块。
首先,利用当前点的坐标得到对应的空间块,其次在之前已编码的LOD层中进行最近邻查找,查找与当前块共面、共线和共点的空间块来得到当前点的N近邻。
当进行共面、共线和共点最近邻查找之后,仍然没有得到当前点的N近邻,则会基于快速查找算法来得到当前点的N近邻,具体算法如下:
如图28所示,当进行属性层间预测时,首先利用当前待编码点的几何坐标得到当前点所对应的莫顿码,其次基于当前点的莫顿码在参考帧中查找到第一个大于当前点莫顿码的参考点(j),其次在[j-searchRange,j+searchRange]范围内进行最近邻查找。
其余具体的更新最近邻的算法和帧间最近邻查找算法一致,在这里不在详述,具体的算法会在帧间最近邻查找算法中提到。
在另一种具体的实现方式中,对于层内最近邻查找,图29示出了一种属性层内最近邻查找的LOD结构示意图。如图29所示,如果层内预测算法开启,即语法元素EnableRefferingSameLoD=1,那么可以允许在层内最近邻查找,如对于LOD1层,当前点P6的最近邻点可以为P1,其他层不允许;如果语法元素EnableRefferingSameLoD=0,那么允许在其他层进行层间查找,如对于LOD1层,当前点P6的最近邻点可以为P4。也就是说,当层内预测算法开启时,会在同一层LOD内,在同层已编码的点集中进行最近邻查找,得到当前点的N近邻(同样进行层间最近邻查找)。
在进行属性层内预测时,会基于快速查找算法进行最近邻查找,具体的算法如如图30所示。其中,当前点用网格表示,假设当前点的莫顿码索引为i,则会在[i+1,i+searchRange]进行最近邻查找。具体的最近邻查找算法与帧间基于块的快速查找算法一致,在这里不再详述。
(ii)帧间最近邻查找:
图28为一种属性帧间预测示意图。如图28所示,当进行属性帧间预测时,首先利用当前待编码点的几何坐标得到当前点所对应的莫顿码,其次基于当前点的莫顿码在参考帧中查找到第一个大于当前点莫顿码的参考点(j),其次在[j-searchRange,j+searchRange]范围内进行最近邻查找。
目前的帧内和帧间进行最近邻查找时,是基于块进行邻域查找的,具体的参见图31。如图31所示,在对当前点(莫顿码索引为i)进行邻域查找时,首先将参考帧中的点按照莫顿码划分成N(N=3)个层,具体的划分算法如下:
·第一层:将假设参考帧的点为numPoints,首先将参考帧中的点每M(M=25=32)个点划分到一个块中;
·第二层:在第一层的基础上,同样按照莫顿码的顺序对第一层的块每M(M=25=32)个块划分到一个块中;
·第三层:在第二层的基础上,同样按照莫顿码的顺序对第一层的块每M(M=25=32)个块划分到一个块中;
最终得到如图31所示的预测结构。
在基于如图31所示的预测结构来进行属性预测,假设当前待编码点的莫顿码索引为i,首先在参考帧中得到第一个大于等于当前点莫顿码的点,索引为j。其次基于j计算得到参考点的块索引,具体计算方式如下:
·第一层:BucketSize_0=25=32;
·第二层:BucketSize_1=25=32×BucketSize_0=1024;
·第三层:BucketSize_2=25=32×BucketSize_1=32768。
假设当前点的预测帧中的参考范围为[j-searchRange,j+searchRange],利用j-searchRange计算得到第三层的起始索引,j+searchRange计算得到第三层的终止索引;其次,首先在第三层的块中判断第二层的一些块是否需要进行最近邻查找,其次到第二层,对于第一层中的每个块判断是否需要进行查找,如果第一层的某些块需要进行最近邻查找,则会对第一层中的一些块中点进行逐点判断来更新最近邻。
下面介绍一下,基于索引计算块的算法,假设当前点对应的莫顿码索引为index,那么对应的第三层块的索引为:
idx_2=index/BucketSize_2          (24)
在得到第三层的块索引idx_2之后,可以利用idx_2得到当前块在第二层对应的块的起始索引和终 止索引:
startIdx1=idx_2×BucketSize_1          (25)
endIdx=idx_2×BucketSize_1+BucketSize_1-1             (26)
同样,基于同样的算法基于第二层块的索引得到第一层块的索引。
在基于块进行最近邻查找时,会首先判断当前块是否需要进行最近邻查找,也就是筛选块的最近邻查找。每个空间块可以通过两个变量进行得到minPos和maxPos,minPos表示的是块的最小值,maxPos表示的是块的最大值。
假设当前点查找的N近邻中最远点的距离为Dist,待编码点的坐标为(x,y,z),当前块表示为(minPos,maxPos),其中minPos为包围盒三个维度上的最小值,maxPos为包围盒三个维度上的最大值,则当前点与包围盒之间的距离D计算如下:
int dx=int(std::max(std::max(minPos[0]-point[0],0),point[0]-maxPos[0]));
int dy=int(std::max(std::max(minPos[1]-point[1],0),point[1]-maxPos[1]));
int dz=int(std::max(std::max(minPos[2]-point[2],0),point[2]-maxPos[2]));
D=dx+dy+dz;
当D小于等于Dist,才会去遍历当前块中的点。
(b)点云属性信息的提升变换编码。
图32为一种提升变换的编码流程示意图。提升变换同样是基于LOD对点云属性进行预测编码。与预测变换的不同之处在于,提升变换首先会对LOD进行高低层的划分,按照LOD生成层的逆序进行预测,并且在预测的过程中引入了更新算子来对低层LOD中点的量化权重进行更新,以提高预测的准确性。这是由于低层LOD中点的属性值会频繁的用于高层LOD中点的属性值预测,低层LOD中的点应具有更大的影响力。
步骤1:分割过程。
分割过程是将完整的LOD层分为低LOD层L(N)和高LOD层H(N)。如果某点云有三层LOD,即(LODl)l=0,1,2,经过分割后,LOD2为高LOD层,记为H(N),(LODl)l=0,1为低LOD层,记为L(N)。
步骤2:预测过程。
高层LOD中的点从低层中选取最近邻点的属性信息作为当前待编码点的属性预测值P(N),预测残差D(N)记为:
D(N)=H(N)-P(N)          (27)
步骤3:更新过程。
对高层LOD中的属性预测残差D(N)进行更新,得到U(N),并利用U(N)对低层LOD中点的属性值进行提升,如下式所示:
L′(N)=L(N)+U(N)             (28)
上述过程将依据LOD从高到低的顺序,不断迭代直至最低层LOD。
由于基于LOD的预测方案使得LOD低层中的点具有更大的影响力,基于提升小波变换的变换方案通过引入量化权重,并且根据预测残差D(N)以及预测点和相邻点之间的距离来更新预测残差,最后利用变换过程中的量化权重来对预测残差进行自适应量化。这里需要注意的是,在解码端可以通过几何重构来确定每个点的量化权重值,因此不要对量化权重进行编码。
(c)区域自适应分层变换。
区域自适应分层变换(RAHT)是一种哈尔小波变换,它可以将点云属性信息从空域变换到频域,进一步减少点云属性之间的相关性。其主要思想是按照八叉树结构,采用自底向上的方式对每一层中的节点分别从X、Y、Z三个维度进行变换(如图34),并迭代直至八叉树的根节点。如图33所示,其基本思想是基于八叉树的层级结构进行小波变换,将属性信息与八叉树节点相关联,对于同一父节点中被占据节点的属性沿着自底向上的方式进行递归变换,对于每一层中的节点分别从X、Y、Z三个维度进行变换,直至变换至八叉树的根节点。在分层变换的过程中,将同层节点变换之后得到的低通/低频(DC)系数传递到下一层的节点继续进行变换,而所有的高通/高频(AC)系数可以通过算术编码器进行编码。
在变换过程中,同一层节点变换之后的DC系数(直流分量)将传递到上一层继续变换,而每一层变换后的AC系数(交流分量)将进行量化编码。下面将介绍主要的变换过程。
图35A为一种RAHT正变换的过程示意图,图35B为一种RAHT逆变换的过程示意图。针对RAHT对应的变换与逆变换过程,假设g′L,2x,y,z和g′L,2x+1,y,z为L层中互为近邻点的两个属性DC系数。经过线性变换后,L-1层的信息为AC系数f′L-1,x,y,z和DC系数g′L-1,x,y,z;然后,f′L-1,x,y,z将不再进行变换,直接进行量化编码,g′L-1,x,y,z将继续寻找近邻进行变换,如果寻找不到,则将其直接传递至L-2层,即RAHT变换仅对存在邻居点的节点有效,没有邻居点的节点将直接传递至上一层。在上述变换过程中,g′L,2x,y,z 和g′L,2x+2,y,z对应的权重(该节点内非空子节点的个数)分别为w′L,2x,y,z和w′L,2x+1,y,z(简写为w′0和w′1),g′L-1,x,y,z的权重为w′L-1,x,y,z,则通用变换公式为:
其中,Tw0,w1为变换矩阵:
变换矩阵会随着各点对应的权重自适应变化更新。上述过程会依据八叉树的划分结构不断迭代更新,直至八叉树的根节点。
简单来讲,在现有的G-PCC属性RAHT区域自适应分层变换帧间预测编码中,通过在属性参数集(Attribute Parameters Set,APS)语法元素中决定是否采用哪种帧间预测编码方案来对属性进行帧间预测,并且通过一个语法元素treeDepth来决定帧间预测编码的启动层数,在之下的RAHT编码层,仅仅采用RAHT帧内预测编码。也就是说,在相关技术中,会定义treeDepth之上的编码层采用哪一种RAHT帧间预测编码方式,对于treeDepth之下的编码层,则采用RAHT帧内预测编码方式。这样的属性编码方案存在两个最为主要的问题:
1、并没有对不同slice层(属性解码层或属性解码层)中不同RAHT编码层AC系数的分部情况进行分析,而是直接在序列集中决定当前序列的属性帧间编码方案;
2、通过在aps中决定帧间预测的层数,由于RAHT编码下层的AC系数帧内相关性相比帧间的AC系数相关性更强,因此往往只在RAHT编码上层启动帧间编码。但是这样的编码方案,并没有完全有效地利用到不同RAHT层AC系数的分部情况,从而导致属性信息的编码效率较低。
基于此,本申请实施例提供了一种解码方法,在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;在第一语法标识信息指示当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定当前层的目标解码模式;根据目标解码模式对当前层中的节点进行属性解码,确定当前层中的节点的属性重建值。
本申请实施例还提供了一种编码方法,在确定当前层的节点允许进行属性预测,以及当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,确定当前层的目标编码模式,并确定第一语法标识信息;第一语法标识信息用于指示当前层是否允许自适应选择帧间预测模式和/或帧内预测模式的情况;根据目标编码模式对当前层中的节点进行属性编码,确定当前层中的节点的属性重建值。
这样,通过为每一层引入对应的属性编码模式,针对每一个层进行属性编码时,在编码端可以自适应地选择每一个片的目标编码模式,并将目标编码模式传递给解码端,使得解码端使用所解析到的目标解码模式来对点云的属性进行属性重建,从而提升了点云属性的编码效率,进而提升了点云的编解码性能。
下面将结合附图对本申请各实施例进行详细说明。
在本申请的一实施例中,参见图36,其示出了本申请实施例提供的一种解码方法的流程示意图。如图36所示,该方法可以包括S101至S103:
S101、在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息。
需要说明的是,在本申请实施例中,该解码方法应用于点云解码器(可简称为“解码器”)。其中,该解码方法具体可以是一种点云属性解码方法,更具体地,可以是一种点云属性RAHT变换预测自适应选择帧间预测或者帧内预测进行解码的方法。
还需要说明的是,在本申请实施例中,这里主要是在属性块头信息参数集(Attribute Brick Header,ABH)中为当前序列中每一层引入对应的属性解码模式,可以针对每一个层(Layer)自适应选择对应的目标解码模式,从而能够提升点云属性的解码效率。
在本申请实施例中,当前层可以为当前视频帧中的其中一层。
在本申请的实施例中,一个视频帧可以理解为一幅图像,举例来说,当前帧可以理解为当前图像,参考帧可以理解为参考图像。
在本申请实施例中,当前层包括至少一个节点。
在本申请实施例中,在解码器侧时,当前层可以称为当前属性解码层、当前解码层、当前片等等。本申请实施例对此不作任何限定。
在本申请实施例中,当前层为沿着第一方向、第二方向和第三方向做一次上采样得到的解码层。其中,第一方向为z轴方向,第二方向为y轴方向,第三方向为x轴方向。
需要说明的是,本申请实施例对第一方向、第二方向和第三方向的顺序不作任何限定。示例性的,可以为第二方向、第一方向和第三方向,也可以为第三方向、第二方向和第一方向。
在本申请实施例中,当前层不局限于沿着第一方向、第二方向和第三方向做一次上采样得到的一个解码层,当前层也可以为沿着第一方向、第二方向和第三方向做一次上采样得到的多个解码层,当前层也可以为一个解码层中的至少一个节点组成的层。本申请实施例对此不作任何限定。
在本申请实施例中,第一语法标识信息用于表征当前层允许自适应选择帧间预测模式和/或帧内预测模式。
在本申请的一些实施例中,解析码流,确定第一语法标识信息的取值的实现可以包括:
若第一语法标识信息的取值为第一值,则确定当前系数组允许自适应选择帧间预测模式和/或帧内预测模式;当前系数组包括至少一个层,当前层为至少一个层的其中一层;
若第一语法标识信息的取值为第二值,则确定当前系数组不允许自适应选择帧间预测模式和/或帧内预测模式。
需要说明的是,当前系数组包括至少一个层,当前层为至少一个层的其中一层。
在本申请的一些实施例中,解析码流,确定第一语法标识信息的取值的实现,可以包括:
若第一语法标识信息的取值为第一值,则确定当前层允许自适应选择帧间预测模式和/或帧内预测模式;
若第一语法标识信息的取值为第二值,则确定当前层不允许自适应选择帧间预测模式和/或帧内预测模式。
需要说明的是,在本申请实施例中,第一值与第二值不同,而且第一值和第二值可以是参数形式,也可以是数字形式。具体地,第一语法标识信息可以是写入在概述(profile)中的参数,也可以是一个标志(flag)的取值,这里对此不作具体限定。
示例性地,对于第一值和第二值而言,第一值可以设置为1,第二值可以设置为0;或者,第一值可以设置为0,第二值可以设置为1;或者,第一值可以设置为true,第二值可以设置为false;或者,第一值可以设置为false,第二值可以设置为true;但是这里并不作具体限定。
在本申请实施例中,以写入码流中的flag为例,假设第一值设置为1(true),第二值设置为0(false),这时候如果第一语法标识信息的取值为0(false),那么可以确定当前层不允许自适应选择帧间预测模式和/或帧内预测模式,即无需执行本申请实施例所述的解码方法;如果第一语法标识信息的取值为1(true),那么可以确定当前层允许自适应选择帧间预测模式和/或帧内预测模式,即需要执行本申请实施例所述的解码方法。
在本申请实施例中,第一语法标识信息起着开关的作用,即当第一语法标识信息为第一值(比如1或真)时,表示启动本申请实施例所述的解码算法,即执行本申请实施例所述的解码算法;当第一语法标识信息为第二值(比如0或假)时,表示不启动本申请实施例所述的解码算法,即不执行本申请实施例所述的解码算法。
可以理解的是,与相关技术中直接在序列集决定当前序列的属性解码模式的方案相比,本申请实施例中通过设置第一语法标识信息可以使得不同的层自适应的选择帧内预测和/或帧间预测进行对节点进行属性解码,可以充分考虑到不同层的AC系数的分布情况,从而可以提高RAHT的解码效率。
在本申请实施例中,确定当前层的节点允许进行属性预测的实现,可以包括以下两种方式:
方式1、根据第六语法标识信息,确定当前层的节点允许进行属性预测;
在本申请实施例中,对于方式1的实现可以包括:
解析码流,确定第六语法标识信息;第六语法标识信息用于指示当前层的节点是否允许进行属性预测;
若第六语法标识信息的取值为第一值,则确定当前层的节点允许进行属性预测;
若第一语法标识信息的取值为第二值,则确定确定当前层的节点不允许进行属性预测。
可以理解的是,在解码端采用方式1来确定当前层的节点是否允许进行属性预测时,仅需要通过解析码流,根据第六语法标识信息的取值来判断即可,这样,可以避免解码端需要重复判断的流程,可以简化解码流程,从而提高解码效率。
方式2、根据当前层的相邻节点数目,确定当前层的节点允许进行属性预测。
在本申请实施例中,对于方式2的实现可以包括:
确定当前层的相邻节点数目;其中,相邻节点包括邻域节点数目以及父节点邻域节点数目;
在相邻节点数目大于或等于预设阈值的情况下,确定当前层的节点允许进行属性预测;
在相邻节点数目小于预设阈值的情况下,确定当前层的节点不允许进行属性预测。
可以理解的是,在本申请实施例中,在解码端采用方式2时,解码端和编码端采用相同的流程来确 定来确定当前层的节点是否允许进行属性预测,这样,编码端不需要向解码端传送指示当前层的节点是否允许进行属性预测的码字,解码端也不需要解析相应的码字,在一定程度上也可以提高解码效率。
对于方式1和方式2的选择,本申请实施例对此不作任何限定,具体可以根据实际应用场景进行选择。
S102、在第一语法标识信息指示当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定当前层的目标解码模式。
在本申请实施例中,在第一语法标识信息指示当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解码器通过解析码流,确定当前层对应的目标解码模式。
在本申请实施例中,目标解码模式可以表示为attr_code_mode[i];其中,i为当前层的索引值。
需要说明的是,只有在当前层满足允许进行属性预测、允许进行帧间预测和允许进行帧内预测这三个条件的情况下,才会赋予索引值i。
示例性的,假设当前层的索引(i)为2,若当前层不满足允许进行属性预测、允许进行帧间预测和允许进行帧内预测这三个条件,则解码器直接跳过当前层,直接进行下一层的属性解码,并将索引2赋予下一层;若当前层满足允许进行属性预测、允许进行帧间预测和允许进行帧内预测这三个条件,则解码器将当前层进行属性解码之后,将索引值进行加1操作(i++),得到更新后的索引值(3),并将索引值3传递给下一层。
在本申请的一些实施例中,S102中解析码流,确定当前层的目标解码模式的实现,可以包括S1021至S1023:
S1021、解码码流,确定属性块头信息参数集。
在本申请实施例中,属性块头信息参数集(Attribute Brick Header,ABH)可以包括至少一个层各自对应的目标解码模式。
S1022、从属性块头信息参数集中,确定第二语法标识信息。
在本申请实施例中,解码器在确定属性块头信息参数集之后,从属性块头信息参数集中,确定当前层对应的第二语法标识信息。其中,第二语法标识信息用于指示当前层的目标解码模式。
S1023、根据第二语法标识信息,确定当前层的目标解码模式。
在本申请实施例中,解码器根据当前层对应的第二语法标识信息的取值,确定当前层的目标解码模式。
需要说明的是,在本申请实施例中,第二语法标识信息的取值可以是参数形式,也可以是数字形式。本申请实施例对此不作任何限定。
在本申请的一些实施例中,目标解码模式包括区域自适应分层帧内变换模式、区域自适应分层帧间变换模式和区域自适应分层结合变换模式;其中,
区域自适应分层帧内变换模式表征采用帧内预测模式对当前层的节点进行属性预测变换解码;区域自适应分层帧间变换模式表征采用帧间预测模式对当前层的节点进行属性预测变换解码;区域自适应分层结合变换模式表征采用帧内预测模式结合帧间预测模式对当前层的节点进行属性预测变换解码。
在本申请的一些实施例中,区域自适应分层帧间变换模式包括第一区域自适应分层帧间变换模式和第二区域自适应分层帧间变换模式;区域自适应分层结合变换模式包括第一区域自适应分层结合变换模式、第二区域自适应分层结合变换模式和第三区域自适应分层结合变换模式;其中,
第一区域自适应分层帧间变换模式表征利用节点的几何信息确定同位预测节点的方式对当前层的节点进行属性预测变换解码;
第二区域自适应分层结合变换模式表征利用参考帧的缓存确定同位预测节点的方式对当前层的节点进行属性预测变换解码;
第一区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和第一区域自适应分层帧间变换模式对当前层的节点进行属性预测变换解码;
第二区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和第二区域自适应分层帧间变换模式对当前层的节点进行属性预测变换解码;
第三区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式、第一区域自适应分层帧间变换模式和第二区域自适应分层帧间变换模式对当前层的节点进行属性预测变换解码。
下面对区域自适应分层帧内变换模式、第一区域自适应分层帧间变换模式和第二区域自适应分层帧间变换模式进行介绍:
1、区域自适应分层帧内变换模式:
在一种具体的实现方式中,针对区域自适应分层帧内预测变换编解码,可以基于RAHT变换编解码的基础上进行预测。如图33所示,RAHT属性变换基于八叉树层级的顺序,由体素级别不断进行变 换直至得到根节点,从而完成整个属性的分层变换编解码。在预测变换编解码中,同样基于八叉树的层级顺序进行属性预测变换编解码,但是是从根节点不断进行变换直至到体素级别。在每一次RAHT属性变换的过程中,是基于2×2×2的块进行属性预测变换编解码。具体的如图37所示。如图37所示,可以看到网格填充块为当前待编解码块,斜线填充块为与当前待编解码块共面和共线的一些邻域块。其中,当前块的属性通过如下方式进行归一化处理:
Anode=∑p∈node attribute(p);
wnode=∑p∈node 1={p∈node};
anode=Anode/wnode
首先,可以通过当前块中包含点的属性得到当前块的属性,即:Anode。通过对当前块中包含点属性进行简单的相加,其次利用当前块的属性与的当前块中点的个数进行归一化处理得到当前块属性的均值anode。利用当前块属性的均值进行属性变换编解码。具体编解码过程参见图38。
如图38所示,这里示出了RAHT属性预测变换编解码的整体流程。其中,(a)为当前块以及共面和共线的一些邻域块,(b)为经过归一化处理后的块,(c)为经过上采样后的块,(d)为当前块的属性,(e)为通过利用当前块的邻域属性进行线性加权拟合得到预测块的属性,最终将对两者分别进行属性变换,得到DC和AC系数,对AC系数进行预测编解码。
其中,当前块的预测属性可以通过利用如图39所示进行线性拟合得到。如图39所示,首先得到当前块的19个邻域块,其次利用邻域块与当前块的每个子块之间的空间几何距离对每个子块的属性进行线性加权预测,最终利用线性加权得到的预测块属性进行变换。具体的属性变换如图40所示。
在图40中,(d)表示属性原始值,对应的属性变换系数如下:
(e)表示属性预测值,对应的属性变换系数如下:
根据属性原始值与属性预测值进行减法运算,可以得到预测残差如下:
2、第一区域自适应分层帧间变换模式:
在另一种具体的实现方式中,针对第一区域自适应分层帧间变换模式,第一区域自适应分层帧间变换模式也称为区域自适应分层帧间预测变换编解码方案一,在G-PCC属性帧间预测中,类似帧内预测编解码的过程。首先,基于几何信息构建RAHT属性变换编解码结构,即:由体素级别不断进行变换直至得到根节点,从而完成整个属性的分层变换编解码。按照这样的方式,构建得到帧内编解码结构和帧间属性编解码结构,具体参见图41。
如图41所示,首先利用当前待编解码节点的几何信息在参考帧中得到待编解码节点的同位预测节点,其次利用参考节点的几何信息和属性信息得到当前待编解码节点的预测属性。
其中,根据如下两种不同的方式得到当前待编解码节点的属性预测值:
①当前节点的帧间预测节点有效:即同位节点存在,则将预测节点的属性直接作为当前待编解码节点的属性预测值;
②当前节点的帧间预测节点无效:即同位节点不存在,则利用帧内相邻节点的属性预测值作为待编解码节点的属性预测值。
最终,利用得到的属性预测值来对当前待编解码节点的属性进行预测。从而完成整个属性的预测编解码。
3、第二区域自适应分层帧间变换模式:
在另一种具体的实现方式中,针对第二区域自适应分层帧间变换模式,第二区域自适应分层帧间变换模式也称为区域自适应分层帧间预测变换编解码方案二,与第一区域自适应分层帧间变换模式不同的是,如果启动第二区域自适应分层帧间变换模式,则首先会基于当前待编解码节点的几何信息构建RAHT属性变换编解码结构,即由体素级别不断进行节点合并,直至得到整个RAHT变换树的根节点, 从而完成整个属性的变换编解码分层结构。其次,在根据RAHT变换结构,由根节点进行划分,得到每个节点的N个子节点(N小于等于8),在帧间预测编解码方案二中,会首先利用RAHT变换对N个子节点的属性进行独立正交变换,得到DC系数(直流分量)和AC系数(交流分量),其次按照以下方式来对N个子节点的AC系数进行属性帧间预测:
①当前节点的帧间预测节点有效:即同位节点存在,则将预测节点的属性直接作为当前待编解码节点的属性预测值;
②前节点可以在参考帧的缓存中查找到与当前节点位置完全相同的节点:即同位节点存在,则会将同位节点中包含的M个子节点的AC系数直接作为当前节点N个子节点的AC系数属性预测值。
a)、如果预测节点的AC系数不为零:则将预测节点的AC系数直接作为预测值;
b)、如果预测节点的AC系数为零,则会将帧内预测对应子节点的AC系数作为预测值;
③当前节点的帧间预测节点无效:即同位节点不存在,则利用帧内相邻节点的属性预测值作为待编解码节点的属性预测值。
在本申请的一些实施例中,S1023中根据第二语法标识信息,确定当前层的目标解码模式的实现,可以包括:根据第二语法标识信息的取值,确定当前层采用帧间预测模式和/或帧内预测模式的目标解码模式。
在本申请的一些实施例中,根据第二语法标识信息的取值,确定当前层采用帧间预测模式和/或帧内预测模式的目标解码模式的实现,可以包括:
若第二语法标识信息的取值为第三值,则确定当前层的目标解码模式为区域自适应分层帧内变换模式;
若第二语法标识信息的取值为第四值,则确定当前层的目标解码模式为第一区域自适应分层帧间变换模式;
若第二语法标识信息的取值为第五值,则确定当前层的目标解码模式为第二区域自适应分层帧间变换模式;
若第二语法标识信息的取值为第六值,则确定当前层的目标解码模式为第一区域自适应分层结合变换模式;
若第二语法标识信息的取值为第七值,则确定当前层的目标解码模式为第二区域自适应分层结合变换模式;
若第二语法标识信息的取值为第八值,则确定当前层的目标解码模式为第三区域自适应分层结合变换模式。
在本申请实施例中,第三值、第四值、第五值、第六值、第七值和第八值不同,需要说明的是,第三值、第四值、第五值、第六值、第七值和第八值可以是参数形式,也可以是数字形式。示例性的,第三值为0,第四值为1、第五值为2、第六值为3、第七值为4、第八值为5。本申请实施例对第三值、第四值、第五值、第六值、第七值和第八值的设置不作任何限制。
S103、根据目标解码模式对当前层中的节点进行属性解码,确定当前层中的节点的属性重建值。
需要说明的是,在本申请实施例中,确定出当前层对应的目标解码模式之后,解码器可以根据目标解码模式对当前层中的节点进行属性解码,进而确定出当前层中的节点的属性重建值。
在本申请实施例中,若目标解码模式为区域自适应分层帧内变换模式,则解码器根据区域自适应分层帧内变换模式对当前层中的节点进行属性解码,进而确定出当前层中的节点的属性重建值;若目标解码模式为第一区域自适应分层帧间变换模式,则解码器根据第一区域自适应分层帧间变换模式对当前层中的节点进行属性解码,进而确定出当前层中的节点的属性重建值;若目标解码模式为第二区域自适应分层帧间变换模式,则解码器根据第二区域自适应分层帧间变换模式对当前层中的节点进行属性解码,进而确定出当前层中的节点的属性重建值;若目标解码模式为第一区域自适应分层结合变换模式,则解码器根据第一区域自适应分层结合变换模式对当前层中的节点进行属性解码,进而确定出当前层中的节点的属性重建值;若目标解码模式为第二区域自适应分层结合变换模式,则解码器根据第二区域自适应分层结合变换模式对当前层中的节点进行属性解码,进而确定出当前层中的节点的属性重建值;若目标解码模式为第三区域自适应分层结合变换模式,则解码器根据第三区域自适应分层结合变换模式对当前层中的节点进行属性解码,进而确定出当前层中的节点的属性重建值。
可以理解的是,在本申请实施例中,首先,在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;然后,在所述第一语法标识信息指示所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定所述当前层的目标解码模式;最后,解码端使用所解析到的目标解码模式来对点云的属性进行属性重建,从而提升了点云属性的解码效率,进而提升了 点云的解码性能。
在本申请的一些实施例中,解码方法还包括:
解析码流,确定第三语法标识信息的取值;第三语法标识信息用于指示当前层所在的当前序列中所包含的层的数量;
获取当前层的索引值,若当前层的索引值为大于等于第九值且小于层的数量,则执行解析码流,确定当前层的目标解码模式的步骤;
若当前层的索引值为大于层的数量,则不执行解析码流,确定当前层的目标解码模式的步骤。
在本申请实施例中,当前层的索引值的确定包括以下两种方式:
方式1、根据第七语法标识信息,确定当前层的索引值;
在本申请实施例中,对于方式1的实现可以包括:
解析码流,确定第七语法标识信息的取值;第七语法标识信息用于指示当前层的索引值。
示例性的,若第七语法标识信息的取值为1,则当前层的索引值为1。
方式2、解码端和编码端采用相同的流程来确定来确定当前层的索引值,这样,编码端不需要向解码端传送指示当前层的索引值的码字,解码端也不需要解析相应的码字,在一定程度上也可以提高解码效率。
在本申请实施例中,当前序列中所包含的层的数量可以表示为attr_code_mode_cnt。需要说明的是,层的数量即为当前序列中可以自适应选择帧间预测模式和/或帧内预测模式对应的解码层的数量。
示例性的,假设当前序列中包含的解码层共有20个,若当前序列满足允许进行属性预测、允许进行帧间预测和允许进行帧内预测的这三个条件的层为10个,则attr_code_mode_cnt的取值为10。
在本申请实施例中,当前层的索引值可以表示为i,其中,i为大于等于0的整数。
在本申请实施例中,第九值为0。
在本申请实施例中,若当前层的索引值为大于等于第九值且小于层的数量,则执行解析码流,确定当前层的目标解码模式的步骤,可以表示为:
for(i=0;i<attr_code_mode_cnt;i++)
                attr_code_mode[i]
可以理解的是,通过第三语法标识信息来指示当前层所在的当前序列中所包含的层的数量,可以实现对当前序列中解码层的索引值满足预设条件(允许属性预测、允许帧间预测和允许帧内预测)的层进行自适应选择帧间预测和/或帧内预测模式,这样,可充分考虑到不同层之间AC系数的分布情况,进而提高属性信息的解码效率。
在本申请的一些实施例中,S103中根据目标解码模式对当前层中的节点进行属性解码,确定当前层中的节点的属性重建值的实现,可以包括S1031至S1034:
S1031、确定当前层中的节点的属性预测值。
在本申请实施例中,解码器根据当前层中的节点的相邻节点的数目,确定当前层中的节点的属性预测值。
需要说明的是,在本申请实施例中,如果当前层中的节点的属性可以进行预测,那么可以利用当前层中的节点的邻域节点的重建属性以及每个邻域节点距离当前节点的几何距离进行线性拟合,得到当前层中的节点的属性预测值。
还需要说明的是,在本申请实施例中,针对当前层中的节点的属性预测,可以是基于帧内属性预测变换,也可以是基于帧间属性预测变换,这里不作具体限定。
在本申请的一些实施例中,S1031中确定当前层中的节点的属性预测值的实现,可以包括:
确定当前层中的节点的相邻节点;其中,相邻节点包括邻域节点和父节点邻域节点;
根据相邻节点对应的属性重建值以及当前层中的节点与相邻节点之间的几何距离进行线性拟合,确定当前层中的节点的属性预测值。
在一种具体的实现方式中,以图38为例,首先确定当前层中的节点的19个邻域节点,然后利用邻域节点与当前节点的每个节点之间的空间几何距离对每个节点的属性进行线性加权预测,最终根据线性加权得到的预测值确定每个节点的属性预测值。
需要说明的是,在本申请实施例中,区域自适应分层变换模式为一种哈尔小波变换,它可以将点云属性信息从空域变换到频域,进一步减少点云属性之间的相关性。其主要思想是按照八叉树结构,采用自底向上的方式对每一层中的节点分别从x、y、z三个维度进行变换,并迭代直至八叉树的根节点。在这里,其基本思想是基于八叉树的层级结构进行小波变换,将属性信息与八叉树节点相关联,对于同一父节点中被占据节点的属性沿着自底向上的方式进行递归变换,对于每一层中的节点分别从x、y、z三个维度进行变换,直至变换至八叉树的根节点。在分层变换的过程中,将同层节点变换之后得到的第 一系数传递到下一层的节点继续进行变换,而所有的第二系数通过算术解码器进行解码确定。
S1032、根据目标解码模式对当前层中的节点的属性预测值进行正向变换,确定当前层中的节点的第一系数值和第二系数预测值。
在本申请实施例中,正向变换即为RAHT正变换,第一系数值即为DC系数,第二系数预测值即为AC系数预测值。
在本申请的一些实施例中,在当前层的目标解码模式为区域自适应分层结合变换模式的情况下,确定当前层中的节点的第二系数预测值,包括:
根据区域自适应分层结合变换模式对当前层中的节点进行正向变换,确定当前层中的节点的第一中间预测值和第二中间预测值;
将当前层中的节点的第一中间预测值和第二中间预测值进行相加,得到当前层中的节点的第二系数预测值。
在本申请实施例中,第一中间预测值可以表示为w1*predIntraVal,第二中间预测值可以表示为w2*predIntraVal。
在本申请的一些实施例中,根据区域自适应分层结合变换模式对当前层中的节点进行正向变换,确定当前层中的节点的第一中间预测值和第二中间预测值,包括:
确定当前层的第一目标权重和第二目标权重;
采用区域自适应分层帧内变换模式,对当前层中的节点进行正向变换,确定当前层中的节点的第一属性预测值;
采用区域自适应分层帧间变换模式,对当前层中的节点进行正向变换,确定当前层中的节点的第二属性预测值;
将当前层中的节点的第一属性预测值与第一目标权重进行相乘,得到当前层中的节点的第一中间预测值;
将当前层中的节点的第二属性预测值与第二目标权重进行相乘,得到当前层中的节点的第二中间预测值。
在本申请实施例中,假设当前节点的RAHT帧内预测值为predIntraVal(第一属性预测值),帧间预测值为predInterVal(第二属性预测值),则最终的预测值为predVal(第二系数预测值)可以表示为:
predVal=w1*predIntraVal+w2*predIntraVal       (34)
在本申请实施例中,第一目标权重可以表示为w1,第二目标权重可以表示为w2。
在本申请的一些实施例中,确定当前层的第一目标权重和第二目标权重,可以包括:
解析码流,确定权重索引值;
在预设权重表中确定与权重索引对应的目标权重组合;其中,目标权重组合包括第一目标权重和第二目标权重。
在本申请实施例中,权重索引值可以为参数形式,也可以为数字形式,本申请实施例对此不作任何限制。
示例性的,权重索引值可以为数字形式,比如权重索引值为2。
在本申请实施例中,目标权重组合包括第一目标权重和第二目标权重。在另一实施例中,目标权重组合包括第一目标权重、第二目标权重和第三目标权重。
需要说明的是,目标权重组合中所包含目标权重的数量与目标解码模式相关。
示例性的,若目标解码模式为区域自适应分层帧内变换模式,则目标权重组合包括第一目标权重w1;若目标解码模式为第一区域自适应分层帧间变换模式,则目标权重组合包括第二目标权重w2;若目标解码模式为第二区域自适应分层帧间变换模式,则目标权重组合包括第三目标权重w3;若目标解码模式为第一区域自适应分层结合变换模式,则目标权重组合包括第一目标权重w1和第二目标权重w2;若目标解码模式为第二区域自适应分层结合变换模式,则目标权重组合包括第一目标权重w1和第三目标权重w3;若目标解码模式为第三区域自适应分层结合变换模式,则目标权重组合包括第一目标权重w1、第二目标权重w2和第三目标权重w3。
可以理解的是,在前层的目标解码模式为区域自适应分层结合变换模式的情况下,会通过对不同RAHT变换层的帧间预测值以及帧内预测值进行合并,按照不同的权重最终得到最佳的预测值,从而可以进一步提升点云属性RAHT解码效率。
在本申请的一些实施例中,在当前层的目标解码模式为区域自适应分层帧间变换模式或区域自适应分层结合变换模式的情况下,在确定当前层中的节点的第二系数预测值之后,方法还包括:
在当前层中的节点的第二系数预测值为第十值的情况下,采用区域自适应分层帧内变换模式,对当前层中的节点进行正向变换,得到当前层中的节点的中间第二系数预测值;
将中间第二系数预测值作为当前层中的节点的第二系数预测值。
在本申请实施例中,第十值为0。
在一具体的实施例中,对于任何一种预测解码模式,首先判断帧间的属性预测值是否等于零,如果不等于零,则会将当前预测值直接作为当前节点AC系数的预测值,否则才会将帧内预测得到的AC系数作为当前节点的AC系数预测值。
S1033、根据第二系数预测值确定当前层的节点对应的第二系数值;
在本申请实施例中,第二系数值也称为AC系数重建值。
在本申请的一些实施例中,根据第二系数预测值确定当前层的节点对应的第二系数值,包括:
解码码流,确定当前层中的节点的第二系数解码残差值;
对第二系数解码残差值进行反量化处理,得到当前层中的节点的第二系数反量化残差值;
根据当前层中的节点对应的第二系数预测值和第二系数反化残差值,确定当前层中的节点的第二系数值。
还需要说明的是,在本申请实施例中,第一系数可以是指低频系数,也可以称为直流分量(Direct Current,DC)系数;第二系数可以是指高频系数,也可以称为交流分量(Alternating Current,AC)系数。在分层变换的过程中,将同一层节点变换之后得到的DC系数传递到下一层的节点继续进行变换,而每一层变换后的AC系数将进行量化解码,故在解码端需要结合解析得到的第二系数解码残差值才可以进一步确定出当前层中的节点的第二系数值。
S1034、根据目标解码模式对当前层中的节点的第一系数值和第二系数值进行逆向变换,确定当前层中的节点的属性重建值。
在本申请实施例中,假设g′L,2x,y,z和g′L,2x+1,y,z为L层中互为近邻点的两个属性DC系数。经过线性变换后,L-1层的信息为AC系数f′L-1,x,y,z和DC系数g′L-1,x,y,z;然后,f′L-1,x,y,z将不再进行变换,直接进行量化解码,g′L-1,x,y,z将继续寻找近邻进行变换,如果寻找不到,则将其直接传递至L-2层,即RAHT变换仅对存在邻居点的节点有效,没有邻居点的节点将直接传递至上一层。在该变换过程中,g′L,2x,y,z和g′L,2x+2,y,z对应的权重(该节点内非空子节点的个数)分别为w′L,2x,y,z和w′L,2x+1,y,z(简写为w′0和w′1),g′L-1,x,y,z的权重为w′L-1,x,y,z,则通用变换公式为:
其中,Tw0,w1为变换矩阵,变换矩阵会随着各点对应的权重自适应变化更新。RAHT的正向变换(也可称为“RAHT正变换”)如前述的图35A所示。
进一步地,根据所得到的当前片中的点的DC系数和AC系数进行RAHT的逆向变换,可以恢复得到当前片中的点的属性重建值。其中,RAHT的逆向变换(也可称为“RAHT反变换”、“RAHT逆变换”)如前述的图35B所示。
在一种具体的实施例中,以当前层中的当前节点为例,解码端的实施步骤如下:
首先,在确定当前节点进行属性预测时,利用当前节点的邻域节点的重建属性以及每个邻域节点距离当前节点的空间几何距离进行线性拟合,得到当前节点的每个子节点的属性预测值;
其次,利用每个子节点的属性预测值,进行RAHT变换得到相应的DC系数和AC系数,最终利用预测节点的AC系数和码流中解析得到的AC系数,来恢复得到当前节点的AC系数;
再次,利用当前节点的AC系数和DC系数进行RAHT反变换,从而恢复得到当前节点的每个子节点的属性重建值。
最后,不断重复前述步骤的顺序,依次从RAHT变换的根节点不断重复直至到RAHT的叶子结点层的最后一个节点,从而完成当前层的RAHT变换的属性解码。
可以理解的是,通过目标解码模式对当前层中的节点进行属性解码,可以考虑到当前层的AC系数的分布情况,从而提高当前层的解码效率。
在本申请的一些实施例中,解码方法还包括:
确定当前层的相邻节点数目;其中,相邻节点包括邻域节点数目以及父节点邻域节点数目;
在相邻节点数目大于或等于预设阈值的情况下,确定当前层的节点允许进行属性预测。
在本申请实施例中,预设阈值用于确定当前层的节点是否允许进行属性预测。
在本申请实施例中,预设阈值为实现设置好的值。预设阈值可以为解码器和编码器双方实现约定好的值。预设阈值也可以为解码器通过解析码流来确定的。本申请实施例对预设阈值的获取方式不作任何限制。
还需要说明的是,在本申请实施例中,如果当前层的节点的相邻节点数目大于或等于预设阈值,那么可以确定当前层允许进行属性预测,此时继续执行基于当前层的邻域节点的属性信息,确定当前层 的节点的属性预测值的步骤;否则,如果当前层的相邻节点数目小于预设阈值,那么可以确定当前层不允许进行属性预测,此时直接停止当前层的节点的属性预测,可以进行下一层的属性预测。
在本申请的一些实施例中,解码方法还包括:
基于当前层中各个节点的空间位置,确定各个节点的邻域节点;
其中,节点的邻域节点至少包括:与节点共面的邻域节点和与节点共线的邻域节点。
需要说明的是,在本申请实施例中,当前层中各个节点的空间位置息可以是节点的位置信息,具体是三维坐标信息(x,y,z)。
在一种具体的实现方式中,节点的邻域节点可以包括:与节点共面的邻域节点和与节点共线的邻域节点。示例性地,如图37所示,网格填充块可以表示当前节点,那么斜线填充块可以表示与当前节点共面和共线的一些邻域节点。
在本申请的一些实施例中,解码方法还包括:
确定当前层中各个节点的父节点;
基于各个节点的父节点的空间位置,确定各个节点的父节点邻域节点;其中,节点的父节点邻域节点至少包括:与节点的父节点共面的邻域节点和与节点的父节点共线的邻域节点。
在本申请的一些实施例中,确定当前层的相邻节点数目,包括:
对当前层中各个节点的邻域节点进行数量统计,确定当前层的邻域节点数目;
对当前层中各个节点的父节点邻域节点进行数量统计,确定当前层的父节点邻域节点数目;
将邻域节点数目和父节点邻域节点数目进行相加,得到当前层的相邻节点数目。
可以理解地,在G-PCC编解码框架中,RAHT既可以用作变换,又可以用作预测,导致复杂度很高。由于考虑到复杂度偏高的问题,相关技术对当前节点是否允许进行属性预测设置了启动条件,具体是:判断当前层的相邻节点数目是否大于预设阈值。这样,通过设置当前层是否启动属性预测的判断条件,可以在保证复杂度的情况下,降低点云属性解码的内存占用率,而且还提升了点云的解码效率。
在本申请的一些实施例中,在第四语法标识信息与第五语法标识信息均为第一值的情况下,第一语法标识信息为第一值;第四语法标识信息用于指示当前层的节点是否允许进行帧间预测;第五语法标识信息用于指示当前层的节点是否允许进行帧内预测;
在第四语法标识信息与第五语法标识信息任意一个为第二值的情况下,第一语法标识信息为第二值。
在本申请实施例中,只有在第四语法标识信息与第五语法标识信息均为第一值的情况下,第一语法元素信息为第一值,即,在确定当前属性解码的节点允许进行帧间预测和允许进行帧内预测的情况下,第一语法元素信息为第一值。
在本申请的一些实施例中,解码方法还包括:
在第一语法标识信息指示当前层不允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定第四语法标识信息;
在第四语法标识信息指示当前层的节点不允许进行帧间预测的情况下,解析码流,确定第五语法标识信息;
在第五语法标识信息指示当前层的节点允许进行帧内预测的情况下,根据区域自适应分层帧内变换模式对当前层中的节点进行属性解码,确定当前层中的节点的属性重建值。
在本申请实施例中,第四语法标识信息可以表示为!disableAttrInterPred,第五语法标识信息可以表示为raht_prediction_enabled。
示例性的,当!disableAttrInterPred为真,则确定当前层的节点允许进行帧间预测;当!disableAttrInterPred为假,则确定当前层的节点不允许进行帧间预测。
示例性的,当raht_prediction_enabled为真(1或true),则确定当前层的节点允许进行帧内预测;当raht_prediction_enabled为假(0或false),则确定当前层的节点不允许进行帧内预测。
在本申请实施例中,第四语法标识信息和第五语法标识信息可以为高层语法元素,第四语法标识信息和第五语法标识信息可以设置与属性参数集(aps)中。
在本申请实施例中,解码器通过解析码流,确定属性参数集;从属性参数集中确定当前层对应的第四语法标识信息和第五语法标识信息。
在本申请实施例中,只有在确定当前层的节点允许进行属性预测,以及确定当前层的节点允许进行帧间预测和帧内预测的情况下,解码器通过解析码流,确定当前层对应的目标解码模式。也就是说,解码器在判断出当前层对应的第一语法标识信息为真(1或true)的情况下,才会继续执行解析码流,确定当前层的目标属性。
在本申请的一些实施例中,解码方法还包括:在第四语法标识信息指示当前层的节点允许进行帧间预测的情况下,根据区域自适应分层帧间变换模式对当前层中的节点进行属性解码,确定当前层中的节 点的属性重建值。
在本申请的一些实施例中,解码方法还包括:
若第四语法标识信息的取值为第一值,则确定当前层的节点允许进行帧间预测;
若第四语法标识信息的取值为第二值,则确定当前层的节点不允许进行帧间预测。
在本申请的一些实施例中,解码方法还包括:
若第五语法标识信息的取值为第一值,则确定当前层的节点允许进行帧间预测;
若第五语法标识信息的取值为第二值(为假),则确定当前层的节点不允许进行帧间预测。
在本申请的一些实施例中,解码方法还包括:
解析码流,确定第六语法标识信息(attr_coding_type)的取值;第六语法标识信息用于指示当前层中的节点采用区域自适应分层帧间变换模式。
在本申请实施例中,第六语法标识信息可以表示为attr_coding_type。
在本申请实施例中,提供了一种解码方法,应用于解码器,首先,在确定当前层的节点允许进行属性预测的情况下,解码器解析码流,确定第一语法标识信息;然后,在第一语法标识信息指示当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解码器解析码流,确定当前层的目标解码模式;最后,解码器使用所解析到的目标解码模式来对点云的属性进行属性重建,从而提升了点云属性的解码效率,进而提升了点云的解码性能。
在本申请的另一实施例中,参见图42,其示出了本申请实施例提供的一种编码方法的流程示意图。如图42所示,该方法可以包括S301至S302:
S301、在确定当前层的节点允许进行属性预测,以及当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,确定当前层的目标编码模式,并确定第一语法标识信息;第一语法标识信息用于指示当前层是否允许自适应选择帧间预测模式和/或帧内预测模式的情况。
需要说明的是,在本申请实施例中,该编码方法应用于点云编码器(可简称为“编码器”)。其中,该编码方法具体可以是一种点云属性编码方法,更具体地,可以是一种点云属性RAHT变换预测自适应选择帧间预测或者帧内预测进行编码的方法。
还需要说明的是,在本申请实施例中,这里主要是在属性块头信息参数集(ABH)中为当前序列中每一层引入对应的目标编码模式,可以针对每一个层(Layer)自适应选择对应的目标编码模式,从而能够提升点云属性的编码效率。
在本申请实施例中,当前层可以为当前视频帧中的其中一层。
在本申请实施例中,当前层包括至少一个节点。
在本申请实施例中,在编码器侧时,当前层可以称为当前属性编码层、当前编码层、当前片等等。本申请实施例对此不作任何限定。
在本申请实施例中,当前层为沿着第一方向、第二方向和第三方向做一次下采样得到的编码层。其中,第一方向为z轴方向,第二方向为y轴方向,第三方向为x轴方向。
需要说明的是,本申请实施例对第一方向、第二方向和第三方向的顺序不作任何限定。示例性的,可以为第二方向、第一方向和第三方向,也可以为第三方向、第二方向和第一方向。
在本申请实施例中,当前层不局限于沿着第一方向、第二方向和第三方向做一次下采样得到的一个编码层,当前层也可以为沿着第一方向、第二方向和第三方向做一次下采样得到的多个编码层,当前层也可以为一个编码层中的至少一个节点组成的层。本申请实施例对此不作任何限定。
在本申请实施例中,第一语法标识信息用于表征当前层允许自适应选择帧间预测模式和/或帧内预测模式。
在本申请的一些实施例中,编码方法还包括:确定第一语法标识信息。
在本申请的一些实施例中,确定第一语法标识信息的实现可以包括:
若确定当前系数组允许自适应选择帧间预测模式和/或帧内预测模式,则确定第一语法标识信息的取值为第一值;当前系数组包括至少一个层,当前层为至少一个层的其中一层;
若确定当前系数组不允许自适应选择帧间预测模式和/或帧内预测模式,则确定第一语法标识信息的取值为第二值。
在本申请的一些实施例中,确定第一语法标识信息的实现还可以包括:
若确定当前层允许自适应选择帧间预测模式和/或帧内预测模式,则确定第一语法标识信息的取值为第一值;
若确定当前层不允许自适应选择帧间预测模式和/或帧内预测模式,则确定第一语法标识信息的取值为第二值。
需要说明的是,在本申请实施例中,第一值与第二值不同,而且第一值和第二值可以是参数形式,也可以是数字形式。具体地,第一语法标识信息可以是写入在概述(profile)中的参数,也可以是一个标志(flag)的取值,这里对此不作具体限定。
示例性地,对于第一值和第二值而言,第一值可以设置为1,第二值可以设置为0;或者,第一值可以设置为0,第二值可以设置为1;或者,第一值可以设置为true,第二值可以设置为false;或者,第一值可以设置为false,第二值可以设置为true;但是这里并不作具体限定。
在本申请实施例中,以写入码流中的flag为例,假设第一值设置为1(true),第二值设置为0(false),这时候如果第一语法标识信息的取值为0(false),那么可以确定当前层不允许自适应选择帧间预测模式和/或帧内预测模式,即无需执行本申请实施例的编码方法;如果第一语法标识信息的取值为1(true),那么可以确定当前层允许自适应选择帧间预测模式和/或帧内预测模式,即需要执行本申请实施例的编码方法。
在本申请实施例中,第一语法标识信息起着开关的作用,即当第一语法标识信息为第一值(比如1或真)时,表示启动本申请实施例的编码算法,即执行本申请实施例的编码算法;当第一语法标识信息为第二值(比如0或假)时,表示不启动本申请实施例的编码算法,即不执行本申请实施例的编码算法。
可以理解的是,与相关技术中直接在序列集决定当前序列的属性编码模式的方案相比,本申请实施例中通过设置第一语法标识信息可以使得不同的层自适应的选择帧内预测和/或帧间预测进行对节点进行属性编码,可以充分考虑到不同层的AC系数的分布情况,从而可以提高RAHT的编码效率。
在本申请实施例中,在确定当前层的节点是否允许进行属性预测之后,确定第六语法标识信息的取值;第六语法标识信息用于指示当前层的节点是否允许进行属性预测;将第六语法标识信息进行编码处理,并将所得到的编码比特写入码流。
在本申请实施例中,确定第六语法标识信息的取值可以包括:若确定当前层的节点允许进行属性预测,则将第六语法标识信息的取值设置为第一值;若确定当前层的节点不允许进行属性预测,则将第六语法标识信息的取值设置为第二值。
可以理解的是,编码器将第六语法标识信息进行编码处理,并将所得到的编码比特写入码流之后,后续在解码端,解码端可以从通过解析码流,根据第六语法标识信息的取值来判断即可。
在本申请申请实施例中,编码器和解码器可以通过双方约定好的判断方式来确定当前层的节点是否允许进行属性预测,在一实施例中,确定当前层的节点允许进行属性预测的实现,可以包括:
确定当前层的相邻节点数目;其中,相邻节点包括邻域节点数目以及父节点邻域节点数目;
在相邻节点数目大于或等于预设阈值的情况下,确定当前层的节点允许进行属性预测;
在相邻节点数目小于预设阈值的情况下,确定当前层的节点不允许进行属性预测。
可以理解的是,编码端和解码端利用约定好的判断方式确定当前层的节点是否允许进行属性预测,这样,编码端不需要将第六语法标识信息进行编码后的码字写入码流,可以节省码字,从而提高编码效率。
在本申请实施例中,目标编码模式可以表示为attr_code_mode[i];其中,i为当前层的索引值。
需要说明的是,只有在当前层满足允许进行属性预测、允许进行帧间预测和允许进行帧内预测这三个条件的情况下,才会赋予索引值i。
示例性的,假设当前层的索引(i)为2,若当前层不满足允许进行属性预测、允许进行帧间预测和允许进行帧内预测这三个条件,则编码器直接跳过当前层,直接进行下一层的属性编码,并将索引2赋予下一层;若当前层满足允许进行属性预测、允许进行帧间预测和允许进行帧内预测这三个条件,则编码器将当前层进行属性编码之后,将索引值进行加1操作(i++),得到更新后的索引值(3),并将索引值3传递给下一层。
在本申请的一些实施例中,目标编码模式包括区域自适应分层帧内变换模式、区域自适应分层帧间变换模式和区域自适应分层结合变换模式;其中,
区域自适应分层帧内变换模式表征采用帧内预测模式对当前层的节点进行属性预测变换编码;区域自适应分层帧间变换模式表征采用帧间预测模式对当前层的节点进行属性预测变换编码;区域自适应分层结合变换模式表征采用帧内预测模式结合帧间预测模式对当前层的节点进行属性预测变换编码。
在本申请的一些实施例中,区域自适应分层帧间变换模式包括第一区域自适应分层帧间变换模式和第二区域自适应分层帧间变换模式;区域自适应分层结合变换模式包括第一区域自适应分层结合变换模式、第二区域自适应分层结合变换模式和第三区域自适应分层结合变换模式;其中,
第一区域自适应分层帧间变换模式表征利用节点的几何信息确定同位预测节点的方式对当前层的节点进行属性预测变换编码;
第二区域自适应分层结合变换模式表征利用参考帧的缓存确定同位预测节点的方式对当前层的节 点进行属性预测变换编码;
第一区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和第一区域自适应分层帧间变换模式对当前层的节点进行属性预测变换编码;
第二区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和第二区域自适应分层帧间变换模式对当前层的节点进行属性预测变换编码;
第三区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式、第一区域自适应分层帧间变换模式和第二区域自适应分层帧间变换模式对当前层的节点进行属性预测变换编码。
需要说明的是,有关区域自适应分层帧内变换模式、第一区域自适应分层帧间变换模式、第二区域自适应分层帧间变换模式、第一区域自适应分层结合变换模式、第二区域自适应分层结合变换模式和第三区域自适应分层结合变换模式的相关描述可参见前文中解码端的相关描述,此处不再赘述。
在本申请的一些实施例中,编码方法还包括:
根据目标编码模式,确定第二语法标识信息;
将第二语法标识信息添加至属性块头信息参数集,并对属性块头信息参数集进行编码处理,将所得到的编码比特写入码流。
在本申请实施例中,编码器根据目标编码模式的取值,确定第二语法标识信息。其中,第二语法标识信息用于指示当前层的目标编码模式。
需要说明的是,在本申请实施例中,第二语法标识信息的取值可以是参数形式,也可以是数字形式。本申请实施例对此不作任何限定。
在本申请的一些实施例中,编码方法还包括:
对当前层中的节点的目标编码模式进行编码处理,将所得到的编码比特写入码流。
在本申请实施例中,编码端采用可以采用率失真算法,得到至少一种候选编码模式对应的代价值,然后根据至少一种候选编码模式对应的代价值,确定目标编码模式,并确定第一语法标识信息;第一语法标识信息用于指示当前层是否允许自适应选择帧间预测模式和/或帧内预测模式的情况,随后将第一语法标识信息和第二语法标识信息写入码流。或者,编码端也可以直接将第二语法标识信息写入码流,这样,在后续的解码侧,解码端可以直接解码第二语法标识信息得到目标解码模式。本申请对此不作任何限定。
需要说明的是,在本申请实施例中,对于编码当前层的目标编码模式,该方法还可以包括:将当前层的目标编码模式添加至属性块头信息参数集;对属性块头信息参数集进行编码处理,将所得到的编码比特写入码流。
还需要说明的是,在本申请实施例中,由于当前序列包括至少一个层,这里是将每个层的目标编码模式均添加到属性块头信息参数集中。因此,在一些实施例中,该方法还可以包括:对当前序列进行编码层划分,确定至少一个编码层;将至少一个编码层呢个各自对应的目标编码模式全部添加至属性块头信息参数集。
这样,在本申请实施例中,后续在解码端,解码端可以从属性块头信息参数集中,直接解码出当前片的目标解码模式。
在本申请的一些实施例中,根据目标编码模式,确定第二语法标识信息的实现,可以包括:
根据当前层采用帧间预测模式和/或帧内预测模式的目标解码模式,确定第二语法标识信息的取值。
在本申请的一些实施例中,根据当前层采用帧间预测模式和/或帧内预测模式的目标解码模式,确定第二语法标识信息的取值的实现,可以包括:
若当前层的目标编码模式为区域自适应分层帧内变换模式,则确定第二语法标识信息的取值为第三值;
若当前层的目标编码模式为第一区域自适应分层帧间变换模式,则确定第二语法标识信息的取值为第四值;
若当前层的目标编码模式为第二区域自适应分层帧间变换模式,则确定第二语法标识信息的取值为第五值;
若当前层的目标编码模式为第一区域自适应分层结合变换模式,则确定第二语法标识信息的取值为第六值;
若当前层的目标编码模式为第二区域自适应分层结合变换模式,则确定第二语法标识信息的取值为第七值;
若当前层的目标编码模式为第三区域自适应分层结合变换模式,则确定第二语法标识信息的取值为第八值。
在本申请实施例中,第三值、第四值、第五值、第六值、第七值和第八值不同,需要说明的是,第 三值、第四值、第五值、第六值、第七值和第八值可以是参数形式,也可以是数字形式。示例性的,第三值为0,第四值为1、第五值为2、第六值为3、第七值为4、第八值为5。本申请实施例对第三值、第四值、第五值、第六值、第七值和第八值的设置不作任何限制。
S302、根据目标编码模式对当前层中的节点进行属性编码,确定当前层中的节点的属性重建值。
需要说明的是,在本申请实施例中,确定出当前层对应的目标编码模式之后,编码器可以根据目标编码模式对当前层中的节点进行属性编码,进而确定出当前层中的节点的属性重建值。
在本申请实施例中,若目标编码模式为区域自适应分层帧内变换模式,则编码器根据区域自适应分层帧内变换模式对当前层中的节点进行属性编码,进而确定出当前层中的节点的属性重建值;若目标编码模式为第一区域自适应分层帧间变换模式,则编码器根据第一区域自适应分层帧间变换模式对当前层中的节点进行属性编码,进而确定出当前层中的节点的属性重建值;若目标编码模式为第二区域自适应分层帧间变换模式,则编码器根据第二区域自适应分层帧间变换模式对当前层中的节点进行属性编码,进而确定出当前层中的节点的属性重建值;若目标编码模式为第一区域自适应分层结合变换模式,则编码器根据第一区域自适应分层结合变换模式对当前层中的节点进行属性编码,进而确定出当前层中的节点的属性重建值;若目标编码模式为第二区域自适应分层结合变换模式,则编码器根据第二区域自适应分层结合变换模式对当前层中的节点进行属性编码,进而确定出当前层中的节点的属性重建值;若目标编码模式为第三区域自适应分层结合变换模式,则编码器根据第三区域自适应分层结合变换模式对当前层中的节点进行属性编码,进而确定出当前层中的节点的属性重建值。
可以理解的是,在本申请实施例中,首先,编码器在确定当前层的节点允许进行属性预测的情况下,确定第一语法标识信息;在第一语法标识信息指示当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,编码器确定当前层的目标编码模式;根据目标编码模式对当前层中的节点进行属性编码,确定当前层中的节点的属性重建值,从而提升了点云属性的编码效率,进而提升了点云的编码性能。
在本申请的一些实施例中,编码方法还包括:
确定第三语法标识信息的取值;第三语法标识信息用于指示当前层所在的当前序列中所包含层的数量;
获取当前层的索引值,若当前层的索引值为大于等于第九值且小于层的数量,则执行确定当前层的目标编码模式的步骤;
若当前层的索引值为大于层的数量,则不执行确定当前层的目标编码模式的步骤;
将第三语法标识信息进行编码处理,将所得到的编码比特写入码流。
在本申请实施例中,当前序列中所包含的层的数量可以表示为attr_code_mode_cnt。需要说明的是,层的数量即为当前序列中可以自适应选择帧间预测模式和/或帧内预测模式对应的编码层的数量。
示例性的,假设当前序列中包含的编码层共有20个,若当前序列满足允许进行属性预测、允许进行帧间预测和允许进行帧内预测的这三个条件的层为10个,则attr_code_mode_cnt的取值为10。
在本申请实施例中,当前层的索引值可以表示为i,其中,i为大于等于0的整数。
在本申请实施例中,第九值为0。
在本申请实施例中,若当前层的索引值为大于等于第九值且小于层的数量,则执行解析码流,确定当前层的目标编码模式的步骤,可以表示为:
for(i=0;i<attr_code_mode_cnt;i++)
            attr_code_mode[i]
可以理解的是,通过第三语法标识信息来指示当前层所在的当前序列中所包含的层的数量,可以实现对当前序列中编码层的索引值满足预设条件(允许属性预测、允许帧间预测和允许帧内预测)的层进行自适应选择帧间预测和/或帧内预测模式,这样,可充分考虑到不同层之间AC系数的分布情况,进而提高属性信息的编码效率。
在本申请的一些实施例中,S301中确定当前层的目标编码模式中的实现,可以包括S3011至S3013:
S3011、基于至少一种候选编码模式对当前层进行属性编码,确定至少一种候选编码模式各自的预编码结果。
在本申请的一些实施例中,至少一种候选编码模式包括以下至少一种:区域自适应分层帧内变换模式、第一区域自适应分层帧间变换模式、第二区域自适应分层帧间变换模式、第一区域自适应分层结合变换模式、第二区域自适应分层结合变换模式、第三区域自适应分层结合变换模式。
S3012、根据至少一种候选编码模式各自的预编码结果进行代价计算,确定至少一种候选编码模式各自的代价值;
还需要说明的是,在本申请实施例中,对至少一种候选编码模式分别进行代价计算,这里的代价可以是指失真值,也可以是指率失真代价值,或者还可以是指其他代价值,在此不作具体限定。
在本申请的一些实施例中,S3012的实现可以包括:根据至少一种候选编码模式各自的预编码结果进行率失真代价计算,确定至少一种候选编码模式各自的代价值。
示例性地,以率失真代价为例,在一些实施例中,根据至少一种候选编码模式各自的预编码结果进行代价计算,确定至少一种候选编码模式各自的代价值,可以包括:根据至少一种候选编码模式各自的预编码结果进行率失真代价计算,确定至少一种候选编码模式各自的代价值。
在本申请实施例中,率失真优化算法中,首先计算每种候选编码模式的重建属性与原始属性的失真D,其次得到每种候选编码模式所需要编码的码流R,那么率失真代价计算如下:
J=D+λ×R             (36)
其中,J表示率失真代价值,R表示候选编码模式所需要编码的码流,λ可以通过属性量化参数进行计算得到,目前的λ计算方式如下:
其中,QP表示量化参数,N可以根据反射率和颜色设定为不同的数值。
S3013、根据至少一种候选编码模式各自的代价值,从至少一种候选编码模式中确定当前层的目标编码模式。
在本申请的一些实施例中,S3023的实现可以包括:
从至少一种候选编码模式各自的代价值中确定最小代价值;
将最小代价值对应的候选编码模式确定为当前层的目标编码模式。
示例性的,假设至少一种候选编码模式包括:第二区域自适应分层结合变换模式和第三区域自适应分层结合变换模式,则针对第二区域自适应分层结合变换模式和第三区域自适应分层结合变换模式分别进行代价计算,确定第二区域自适应分层结合变换模式的第一代价值以及第三区域自适应分层结合变换模式的第二代价值;根据第一代价值和第二代价值,确定当前层的目标编码模式。
在本申请实施例中,根据第一代价值和第二代价值,确定当前层的目标编码模式,具体可以是:若第一代价值小于第二代价值,则将第二区域自适应分层结合变换模式确定为当前层的目标编码模式;或者,若第一代价值大于第二代价值,则将第三区域自适应分层结合变换模式确定为当前层的目标编码模式。
另外,在本申请实施例中,如果第一代价值等于第二代价值,那么可以将第二区域自适应分层结合变换模式确定为当前层的目标编码模式,或者也可以将第三区域自适应分层结合变换模式确定为当前层的目标编码模式,这里不作具体限定。
在本申请的一些实施例中,S302中根据目标编码模式对当前层中的节点进行属性编码,确定当前层中的节点的属性重建值的实现,可以包括S3021至S3024:
S3021、确定当前层中的节点的属性预测值。
在本申请实施例中,编码器根据当前层中的节点的相邻节点的数目,确定当前层中的节点的属性预测值。
需要说明的是,在本申请实施例中,如果当前层中的节点的属性可以进行预测,那么可以利用当前层中的节点的邻域节点的重建属性以及每个邻域节点距离当前节点的几何距离进行线性拟合,得到当前层中的节点的属性预测值。
还需要说明的是,在本申请实施例中,针对当前层中的节点的属性预测,可以是基于帧内属性预测变换,也可以是基于帧间属性预测变换,这里不作具体限定。
在本申请的一些实施例中,确定当前层中的节点的属性预测值的实现,可以包括:
确定当前层中的节点的相邻节点;其中,相邻节点包括邻域节点和父节点邻域节点;
根据相邻节点对应的属性重建值以及当前层中的节点与相邻节点之间的几何距离进行线性拟合,确定当前层中的节点的属性预测值。
在一种具体的实现方式中,首先确定当前层中的节点的19个邻域节点,然后利用邻域节点与当前节点的每个节点之间的空间几何距离对每个节点的属性进行线性加权预测,最终根据线性加权得到的预测值确定每个节点的属性预测值。
需要说明的是,在本申请实施例中,区域自适应分层变换模式为一种哈尔小波变换,它可以将点云属性信息从空域变换到频域,进一步减少点云属性之间的相关性。其主要思想是按照八叉树结构,采用自底向上的方式对每一层中的节点分别从x、y、z三个维度进行变换,并迭代直至八叉树的根节点。在这里,其基本思想是基于八叉树的层级结构进行小波变换,将属性信息与八叉树节点相关联,对于同一父节点中被占据节点的属性沿着自底向上的方式进行递归变换,对于每一层中的节点分别从x、y、z三个维度进行变换,直至变换至八叉树的根节点。在分层变换的过程中,将同层节点变换之后得到的第 一系数传递到下一层的节点继续进行变换,而所有的第二系数通过算术编码器进行编码确定。
S3022、根据目标编码模式对当前层中的节点的属性预测值进行正向变换,确定当前层中的节点的第一系数值和第二系数预测值。
在本申请实施例中,正向变换即为RAHT正变换,第一系数值即为DC系数,第二系数预测值即为AC系数预测值。
在本申请的一些实施例中,在当前层的目标编码模式为区域自适应分层结合变换模式的情况下,确定当前层中的节点的第二系数预测值,包括:
根据区域自适应分层结合变换模式对当前层中的节点进行正向变换,确定当前层中的节点的第一中间预测值和第二中间预测值;
将当前层中的节点的第一中间预测值和第二中间预测值进行相加,得到当前层中的节点的第二系数预测值。
在本申请实施例中,第一中间预测值可以表示为w1*predIntraVal,第二中间预测值可以表示为w2*predIntraVal。
在本申请的一些实施例中,根据区域自适应分层结合变换模式对当前层中的节点进行正向变换,确定当前层中的节点的第一中间预测值和第二中间预测值的实现,可以包括:
在预设权重表中确定当前层对应的目标权重组合;其中,目标权重组合包括第一目标权重和第二目标权重;
采用区域自适应分层帧内变换模式,对当前层中的节点进行正向变换,确定当前层中的节点的第一属性预测值;
采用区域自适应分层帧间变换模式,对当前层中的节点进行正向变换,确定当前层中的节点的第二属性预测值;
将当前层中的节点的第一属性预测值与第一目标权重进行相乘,得到当前层中的节点的第一中间预测值;
将当前层中的节点的第二属性预测值与第二目标权重进行相乘,得到当前层中的节点的第二中间预测值。
在本申请实施例中,权重索引值可以为参数形式,也可以为数字形式,本申请实施例对此不作任何限制。
示例性的,权重索引值可以为数字形式,比如权重索引值为2。
在本申请实施例中,目标权重组合包括第一目标权重和第二目标权重。在另一实施例中,目标权重组合包括第一目标权重、第二目标权重和第三目标权重。
需要说明的是,目标权重组合中所包含目标权重的数量与目标编码模式相关。
示例性的,若目标编码模式为区域自适应分层帧内变换模式,则目标权重组合包括第一目标权重w1;若目标编码模式为第一区域自适应分层帧间变换模式,则目标权重组合包括第二目标权重w2;若目标编码模式为第二区域自适应分层帧间变换模式,则目标权重组合包括第三目标权重w3;若目标编码模式为第一区域自适应分层结合变换模式,则目标权重组合包括第一目标权重w1和第二目标权重w2;若目标编码模式为第二区域自适应分层结合变换模式,则目标权重组合包括第一目标权重w1和第三目标权重w3;若目标编码模式为第三区域自适应分层结合变换模式,则目标权重组合包括第一目标权重w1、第二目标权重w2和第三目标权重w3。
可以理解的是,在当前层的目标编码模式为区域自适应分层结合变换模式的情况下,会通过对不同RAHT变换层的帧间预测值以及帧内预测值进行合并,按照不同的权重最终得到最佳的预测值,从而可以进一步提升点云属性RAHT编码效率。
在本申请的一些实施例中,将目标权重组合在预设权重表中对应的权重索引值进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,在当前层的目标编码模式为区域自适应分层帧间变换模式或区域自适应分层结合变换模式的情况下,在确定当前层中的节点的第二系数预测值之后,方法还包括:
在当前层中的节点的第二系数预测值为第十值的情况下,根据区域自适应分层帧内变换模式对当前层中的节点进行正向变换,得到当前层中的节点的中间第二系数预测值;
将中间第二系数预测值作为当前层中的节点的第二系数预测值。
还需要说明的是,在本申请实施例中,第一系数可以是指低频系数,也可以称为直流分量(Direct Current,DC)系数;第二系数可以是指高频系数,也可以称为交流分量(Alternating Current,AC)系数。在分层变换的过程中,将同一层节点变换之后得到的DC系数传递到下一层的节点继续进行变换,而每一层变换后的AC系数将进行量化编码。
S3023、根据第二系数预测值确定当前层的节点对应的第二系数值。
在本申请的一些实施例中,根据第二系数预测值确定当前层的节点对应的第二系数值的实现,可以包括:
确定当前层中的节点的第二系数编码残差值;
对第二系数编码残差值进行反量化处理,得到当前层中的节点的第二系数反量化残差值;
根据当前层中的节点对应的第二系数预测值和第二系数反化残差值,确定当前层中的节点的第二系数值。
在本申请的一些实施例中,确定当前层中的节点的第二系数编码残差值的实现,可以包括:
确定当前层中的节点的属性原始值;
根据目标编码模式对当前层中的节点的属性原始值进行正向变换,确定当前层中的节点的第一系数值和第二系数原始值;
根据当前层中的点的第二系数原始值和第二系数预测值,确定当前层中的节点的第二系数预测残差值;
对第二系数残差值进行量化处理,得到当前层中的节点的第二系数量化残差值。
在本申请的一些实施例中,编码方法还包括:
对当前层中的节点的第二系数量化残差值进行编码处理,将所得到的编码比特写入码流。
S3024、根据目标编码模式对当前层中的节点的第一系数值和第二系数值进行逆向变换,确定当前层中的节点的属性重建值。
需要说明的是,在本申请实施例中,第一系数可以是指低频系数,也可以称为直流分量(Direct Current,DC)系数;第二系数可以是指高频系数,也可以称为交流分量(Alternating Current,AC)系数。在分层变换的过程中,将同一层节点变换之后得到的DC系数传递到下一层的节点继续进行变换,而每一层变换后的AC系数将进行量化编码,以便后续在解码端能够确定出当前片中的点的第二系数值。
还需要说明的是,在本申请实施例中,假设g′L,2x,y,z和g′L,2x+1,y,z为L层中互为近邻点的两个属性DC系数。经过线性变换后,L-1层的信息为AC系数f′L-1,x,y,z和DC系数g′L-1,x,y,x;然后,f′L-1,x,y,z将不再进行变换,直接进行量化编码,g′L-1,x,y,z将继续寻找近邻进行变换,如果寻找不到,则将其直接传递至L-2层,即RAHT变换仅对存在邻居点的节点有效,没有邻居点的节点将直接传递至上一层。在该变换过程中,g′L,2x,y,z和g′L,2x+2,y,z对应的权重(该节点内非空子节点的个数)分别为w′L,2x,y,z和w′L,2x+1,y,z(简写为w′0和w′1),g′L-1,x,y,z的权重为w′L-1,x,y,z,则通用变换公式为:
其中,Tw0,w1为变换矩阵,变换矩阵会随着各点对应的权重自适应变化更新。RAHT的正向变换(也可称为“RAHT正变换”)如前述的图35A所示。
还需要说明的是,在本申请实施例中,根据所得到的当前片中的点的DC系数和AC系数进行RAHT的逆向变换,可以恢复得到当前片中的点的属性重建值。其中,RAHT的逆向变换(也可称为“RAHT反变换”、“RAHT逆变换”)如前述的图35B所示。
在一种具体的实施例中,以当前层中的当前节点为例,编码端的实施步骤如下:
首先,在确定当前节点进行属性预测时,利用当前节点的邻域节点的重建属性以及每个邻域节点距离当前节点的空间几何距离进行线性拟合,得到当前节点的每个子节点的属性预测值;
其次,利用每个子节点的属性预测值,进行RAHT变换得到相应的DC系数和AC系数。同样的,通过RAHT变换来对当前节点的每个子节点的属性进行变换得到DC系数和AC系数;
再次,利用预测节点得到的AC系数的预测值来对当前节点的AC进行预测,最终对每个子节点的AC预测残差系数进行量化编码。
再次,利用AC预测残差系数的反量化值以及AC系数的预测值,恢复得到当前节点的AC重建系数,最终利用当前节点的AC系数和DC系数进行RAHT反变换,从而恢复得到当前节点每个子节点的属性重建值。
最后,不断重复前述步骤的顺序,依次从RAHT变换的根节点不断重复直至到RAHT的叶子结点层的最后一个节点,从而完成当前层的RAHT变换的属性编码。
也就是说,在本申请实施例中,首先,编码器在确定当前层的节点允许进行属性预测的情况下,确定第一语法标识信息;在第一语法标识信息指示当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,编码器确定当前层的目标编码模式;根据目标编码模式对当前层中的节点进行属性编码,确定当前层中的节点的属性重建值,从而提升了点云属性的编码效率,进而提升了点云的编码性能。
在本申请的一些实施例中,编码方法还包括:
确定第四语法标识信息和第五语法标识信息的取值;第五语法标识信息用于指示当前层的节点是否允许进行帧内预测;
在第四语法标识信息与第五语法标识信息均为第一值的情况下,则确定第一语法标识信息的取值为第一值;
在第四语法标识信息与第五语法标识信息任意一个为第二值的情况下,则确定第一语法标识信息的取值为第二值;
将第一语法标识信息进行编码处理,将所得到的编码比特写入码流。
在本申请实施例中,只有在第四语法标识信息与第五语法标识信息均为第一值的情况下,第一语法元素信息为第一值,即,在确定当前属性编码的节点允许进行帧间预测和允许进行帧内预测的情况下,第一语法元素信息为第一值。
在本申请实施例中,第四语法标识信息可以表示为!disableAttrInterPred,第五语法标识信息可以表示为raht_prediction_enabled。
示例性的,当!disableAttrInterPred为真,则确定当前层的节点允许进行帧间预测;当!disableAttrInterPred为假,则确定当前层的节点不允许进行帧间预测。
示例性的,当raht_prediction_enabled为真(1或true),则确定当前层的节点允许进行帧内预测;当raht_prediction_enabled为假(0或false),则确定当前层的节点不允许进行帧内预测。
在本申请的一些实施例中,编码方法还包括:
若确定当前层的节点允许进行帧间预测,则将第四语法标识信息的取值设置为第一值;
若确定当前层的节点不允许进行帧间预测,则将第四语法标识信息的取值设置为第二值;
将第四语法标识信息进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,编码方法还包括:
若确定当前层的节点允许进行帧内预测,则将第五语法标识信息的取值设置为第一值;
若确定当前层的节点不允许进行帧内预测,则将第五语法标识信息的取值设置为第二值;
将第五语法标识信息进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,编码方法还包括:
确定第六语法标识信息的取值;第六语法标识信息用于指示当前层中的节点采用区域自适应分层帧间变换模式;
将第六语法标识信息进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,编码方法还包括:
确定当前层的相邻节点数目;其中,相邻节点包括邻域节点数目以及父节点邻域节点数目;
在相邻节点数目大于或等于预设阈值的情况下,确定当前层的节点允许进行属性预测。
在本申请的一些实施例中,编码方法还包括:
基于当前层中各个节点的空间位置,确定各个节点的邻域节点;
其中,节点的邻域节点至少包括:与节点共面的邻域节点和与节点共线的邻域节点。
在本申请的一些实施例中,编码方法还包括:
确定当前层中各个节点的父节点;
基于各个节点的父节点的空间位置,确定各个节点的父节点邻域节点;其中,节点的父节点邻域节点至少包括:与节点的父节点共面的邻域节点和与节点的父节点共线的邻域节点。
需要说明的是,在本申请实施例中,当前层中各个节点的空间位置息可以是节点的位置信息,具体是三维坐标信息(x,y,z)。
在一种具体的实现方式中,节点的邻域节点可以包括:与节点共面的邻域节点和与节点共线的邻域节点。示例性地,如图37所示,网格填充块可以表示当前节点,那么斜线填充块可以表示与当前节点共面和共线的一些邻域节点。
在本申请的一些实施例中,编码确定当前层的相邻节点数目,包括:
对当前层中各个节点的邻域节点进行数量统计,确定当前层的邻域节点数目;
对当前层中各个节点的父节点邻域节点进行数量统计,确定当前层的父节点邻域节点数目;
将邻域节点数目和父节点邻域节点数目进行相加,得到当前层的相邻节点数目。
可以理解地,在G-PCC编编码框架中,RAHT既可以用作变换,又可以用作预测,导致复杂度很高。由于考虑到复杂度偏高的问题,相关技术对当前节点是否允许进行属性预测设置了启动条件,具体是:判断当前层的相邻节点数目是否大于预设阈值。这样,通过设置当前层是否启动属性预测的判断条件,可以在保证复杂度的情况下,降低点云属性编码的内存占用率,而且还提升了点云的编码效率。
在本申请实施例中,首先,编码器在确定当前层的节点允许进行属性预测的情况下,确定第一语法 标识信息;在第一语法标识信息指示当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,编码器确定当前层的目标编码模式;根据目标编码模式对当前层中的节点进行属性编码,确定当前层中的节点的属性重建值,从而提升了点云属性的编码效率,进而提升了点云的编码性能。
在本申请的又一实施例中,基于前述实施例所述的编解码方法,这里具体提出一种点云属性RAHT变换预测RDO自适应选择帧间预测或者帧内预测编码方案。在编码端,根据RDO方式进行自适应地选择每个属性编码层的最佳编码模式,其次将最终的最佳编码模式传递给解码端,解码端利用得到的最佳编码模式对点云的属性进行属性重建,从而可以进一步提升点云属性的编码效率。
在本申请实施例中,通过引入一种新的编码方案提升点云属性的编码效率,首先通过结合三种属性预测编码方案:两种帧间预测编码方案、一种帧内预测编码方案,并且在对不同RAHT编码层的AC系数编码之前,在编码端利用率失真优化算法得到当前RAHT编码层的最佳编码模式,即:帧间预测编码方案二+帧内预测编码、帧间预测编码方案一+帧间预测编码方案二+帧内预测编码,最终将当前RAHT编码层的最佳编码模式传递给解码端,解码端利用当前层RAHT的编码模式来自适应地恢复当前层的AC系数,从而来完成整个属性RAHT编码,最终提升RAHT属性编码效率。
在本申请实施例中,如图43所示,首先定义RAHT属性编码层(即当前层),目前的属性RAHT变换编码顺序是从根节点依次进行划分直至划分到体素级别(1x1x1),从而完成整个点云属性的编码和属性重建。在这里,定义每次沿着Z方向、Y方向和X方向做一次下采样得到的层即为一个RAHT变换层,即layer。其次,基于RAHT层,引入率失真优化算法,来对当前层的预测编码方式进行自适应选择,引入两种预测编码模式:1、帧内结合帧间预测编码模式二(即第二区域自适应分层结合变换模式);2、帧内预测模式结合帧间预测编码模式二和帧间预测编码模式一(即第三区域自适应分层结合变换模式)。
进一步地,利用率失真优化算法在编码端利用两种预测模式进行预测编码当前层节点的属性信息,最终利用率失真优化算法得到当前层的最佳编码模式,并且将最佳编码模式传递给解码端,解码端利用解析得到的预测解码模式来对当前待解码层点的属性信息进行重建恢复。其中,率失真优化算法中,首先计算每种预测模式的重建属性与原始属性的失真D,其次得到每种预测模式所需要编码的码流R,则率失真代价计算如如上述的式(36)和式(37)所示。
在本申请实施例中,这里是将每一层的目标编码模式最终添加到ABH参数集中。在一种具体的实施例中,编码端具体算法如下:
步骤1:根据当前层的邻域节点数目以及父节点邻域节点数目自适应决定当前层的节点是否可以采用属性预测;
步骤2:如果当前层的节点可以采用属性预测,并且可以进行属性帧间预测,则引入率失真优化算法对于当前层,通过编码当前层的每个节点,计算得到每种预测编码模式对应的代价,得到最佳的预测编码模式。
步骤3:最终利用最佳的预测编码模式来对当前层节点的属性进行预测编码。
在另一种具体的实施例中,解码端具体算法如下:
步骤1:根据当前层的邻域节点数目以及父节点邻域节点数目自适应决定当前层的节点是否可以采用属性预测;
步骤2:如果当前层的节点可以采用属性预测,并且可以进行属性帧间预测,则节点得到当前层最佳的预测解码模式。
步骤3:最终利用最佳的预测解码模式来对当前层节点的属性进行预测解码。
进一步地,在本申请实施例中,针对aps中的语法元素(Attribute parameter set data unit syntax)的描述如表2所示。
表2


简单来说,本申请实施例通过在对属性进行RAHT预测编码时,在每一个RAHT编码层引入一个预测编码模式(attr_code_mode[i])来自适应地选择帧间预测编码模式二结合帧内预测编码模式或者帧内预测编码结合两种帧间预测编码模式,并且最终将该编码模式传递给解码端,解码端利用编码模式来对点云的属性进行重建。在本方案当中,重点是通过在每个RAHT编码层引入一个编码模式,通过在编码端利用率失真优化选择算法得到最佳的编码模式,其次在解码端利用解码模式来对点云的属性进行重建。目前将每层的编码模式存储在ABH中,在解码端通过ABH来得到RAHT的编码层的解码模式,对于该参数以何种形式进行编码,在这里不受限制。
在本申请实施例中,可以对属性帧间预测模式可以进一步进行修正。具体的,在主方案中,是通过结合现有的三种预测预测编码模式:帧间预测编码模式二结合帧内预测编码模式、帧间预测编码模式二、帧间预测编码模式一以及帧内预测编码模式,在编码端对当前RAHT编码层引入一种编码模式来代表采用哪种预测编码模式来恢复当前RAHT编码层的AC系数。该方案可以将预测编码模式进一步更改为:帧间预测编码模式一、帧内预测编码模式、帧间预测编码模式一,按照主方案同样的方式,来决定当前层的最佳编码模式。解码端同样根据当前层的预测编码模式,来恢复得到当前层的AC系数,从而完成整个RAHT属性编码。
在本申请实施例中,可以对属性预测模式可以进一步进行修正。具体的,在主方案中,对于任何一种预测编码模式,首先判断帧间的属性预测值是否等于零,如果不等于零,则会将当前预测值直接作为当前节点AC系数的预测值,否则才会将帧内预测得到的AC系数作为当前节点的AC系数预测值。本方案中,会通过对不同RAHT变换层的帧间预测值以及帧内预测值进行合并,按照不同的权重最终得到最佳的预测值,从而可以进一步提升点云属性RAHT编码效率,具体的预测编码方案如公式(X)所示。
在本申请实施例中,基于上述实施例对前述实施例的具体实现进行了详细阐述,从中可以看出,根据前述实施例的技术方案,这里提出了一种基于层的率失真优化编码算法,通过在对属性进行帧间RAHT预测时,如果当前待编码层可以进行属性预测,则会首先对当前待编码层引入两种编码模式,其次利用率失真优化算法选取最佳的预测编码模式进行预测编码,从而可以提升点云属性的编码效率,进一步地,将最终最佳的编码模式传递给解码端,解码端利用得到的最佳编码模式来对点云的属性进行属性重建,从而可以进一步提升点云属性的编码效率。示例性地,表3示出了关于属性的编码效率的测试结果。
表3
根据表3可以看出,在引入率失真优化算法之后,对于可以采用帧间属性预测的序列,属性编码BPP降低约3.9%,显著提升点云属性的编码效率。
在本申请的一实施例中,基于前述实施例相同的发明构思,提供一种码流,其中,码流是根据待编码信息进行比特编码生成的;其中,待编码信息包括下述至少一项:第一语法标识信息的取值、第二语法标识信息的取值、第三语法标识信息的取值、第四语法标识信息的取值、第五语法标识信息的取值、当前层中的节点对应的权重索引值、当前层中的节点的第二系数量化残差值;
其中,第一语法标识信息用于指示当前层是否允许自适应选择帧间预测模式和/或帧内预测模式,第二语法标识信息用于指示当前层的目标编码模式,第三语法标识信息的取值用于指示当前层所在的当前序列中所包含层的数量,第四语法标识信息的取值用于指示当前层的节点是否允许进行帧间预测,第五语法标识信息用于指示当前层的节点是否允许进行帧内预测,第六语法标识信息用于指示当前层中的节点采用区域自适应分层帧间变换模式,权重索引值用于指示当前层中的节点对应的目标权重组合在预设权重表中对应的索引值。
在本申请的再一实施例中,基于前述实施例相同的发明构思,参见图44,其示出了本申请实施例提供的一种解码器的组成结构示意图。如图44所示,该解码器1000可以包括第一确定部分1001和解码部分1002;其中,
所述第一确定部分1001,被配置为在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;在所述第一语法标识信息指示所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定所述当前层的目标解码模式;
所述解码部分1002,被配置为根据所述目标解码模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为若所述第一语法标识信息的取值为第一值,则确定当前系数组允许自适应选择帧间预测模式和/或帧内预测模式;所述当前系数组包括至少一个层,所述当前层为所述至少一个层的其中一层;若所述第一语法标识信息的取值为第二值,则确定所述当前系数组不允许自适应选择帧间预测模式和/或帧内预测模式。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为若所述第一语法标识信息的取值为第一值,则确定所述当前层允许自适应选择帧间预测模式和/或帧内预测模式;若所述第一语法标识信息的取值为第二值,则确定所述当前层不允许自适应选择帧间预测模式和/或帧内预测模式。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为解码码流,确定属性块头信息参数集;从所述属性块头信息参数集中,确定第二语法标识信息;根据所述第二语法标识信息,确定所述当前层的所述目标解码模式。
在本申请的一些实施例中,所述目标解码模式包括区域自适应分层帧内变换模式、区域自适应分层帧间变换模式和区域自适应分层结合变换模式;其中,所述区域自适应分层帧内变换模式表征采用帧内预测模式对所述当前层的节点进行属性预测变换解码;所述区域自适应分层帧间变换模式表征采用帧间预测模式对所述当前层的节点进行属性预测变换解码;所述区域自适应分层结合变换模式表征采用帧内预测模式结合帧间预测模式对所述当前层的节点进行属性预测变换解码。
在本申请的一些实施例中,所述区域自适应分层帧间变换模式包括第一区域自适应分层帧间变换模式和第二区域自适应分层帧间变换模式;所述区域自适应分层结合变换模式包括第一区域自适应分层结合变换模式、第二区域自适应分层结合变换模式和第三区域自适应分层结合变换模式;其中,
所述第一区域自适应分层帧间变换模式表征利用节点的几何信息确定同位预测节点的方式对所述当前层的节点进行属性预测变换解码;
所述第二区域自适应分层结合变换模式表征利用参考帧的缓存确定同位预测节点的方式对所述当前层的节点进行属性预测变换解码;
所述第一区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和所述第一区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换解码;
所述第二区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和所述第二区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换解码;
所述第三区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式、所述第一区域自适应分层帧间变换模式和所述第二区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换解码。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为根据所述第二语法标识信息的取值,确定所述当前层采用帧间预测模式和/或帧内预测模式的所述目标解码模式。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为若所述第二语法标识信息的取值为第三值,则确定所述当前层的所述目标解码模式为区域自适应分层帧内变换模式;
若所述第二语法标识信息的取值为第四值,则确定所述当前层的所述目标解码模式为第一区域自适应分层帧间变换模式;
若所述第二语法标识信息的取值为第五值,则确定所述当前层的所述目标解码模式为第二区域自适应分层帧间变换模式;
若所述第二语法标识信息的取值为第六值,则确定所述当前层的所述目标解码模式为第一区域自适应分层结合变换模式;
若所述第二语法标识信息的取值为第七值,则确定所述当前层的所述目标解码模式为第二区域自适应分层结合变换模式;
若所述第二语法标识信息的取值为第八值,则确定所述当前层的所述目标解码模式为第三区域自适应分层结合变换模式。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为解析码流,确定第三语法标识信息的取值;所述第三语法标识信息用于指示所述当前层所在的当前序列中所包含的层的数量;获取所述当前层的索引值,若所述当前层的索引值为大于等于第九值且小于所述层的数量,则执行所述解析码流,确定所述当前层的目标解码模式的步骤;若所述当前层的索引值为大于所述层的数量,则不执行所述解析码流,确定所述当前层的目标解码模式的步骤。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为确定所述当前层中的节点的属性预测值;根据所述目标解码模式对所述当前层中的节点的属性预测值进行正向变换,确定所述当前层中的节点的第一系数值和第二系数预测值;根据所述第二系数预测值确定所述当前层的节点对应的第二系数值;根据所述目标解码模式对所述当前层中的节点的所述第一系数值和所述第二系数值进行逆向变换,确定所述当前层中的节点的属性重建值。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为确定所述当前层中的节点的相邻节点;其中,所述相邻节点包括邻域节点和父节点邻域节点;根据所述相邻节点对应的属性重建值以及所述当前层中的节点与所述相邻节点之间的几何距离进行线性拟合,确定所述当前层中的节点的属性预测值。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为解码码流,确定所述当前层中的节点的第二系数解码残差值;对所述第二系数解码残差值进行反量化处理,得到所述当前层中的节点的第二系数反量化残差值;根据所述当前层中的节点对应的第二系数预测值和所述第二系数反化残差值,确定所述当前层中的节点的第二系数值。
在本申请的一些实施例中,在当前层的目标解码模式为所述区域自适应分层结合变换模式的情况下,所述第一确定部分1001,还被配置为根据所述区域自适应分层结合变换模式对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一中间预测值和第二中间预测值;将所述当前层中的节点的所述第一中间预测值和所述第二中间预测值进行相加,得到所述当前层中的节点的所述第二系数预测值。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为确定所述当前层的第一目标权重和第二目标权重;采用所述区域自适应分层帧内变换模式,对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一属性预测值;采用所述区域自适应分层帧间变换模式,对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第二属性预测值;将所述当前层中的节点的所述第一属性预测值与所述第一目标权重进行相乘,得到所述当前层中的节点的所述第一中间预测值;将所述当前层中的节点的所述第二属性预测值与所述第二目标权重进行相乘,得到所述当前层中的节点的所述第二中间预测值。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为解析码流,确定权重索引值;在预设权重表中确定与所述权重索引对应的目标权重组合;其中,所述目标权重组合包括第一目标权重和第二目标权重。
在本申请的一些实施例中,在当前层的目标解码模式为所述区域自适应分层帧间变换模式或所述区域自适应分层结合变换模式的情况下,所述第一确定部分1001,还被配置为在所述当前层中的节点的第二系数预测值为第十值的情况下,采用所述区域自适应分层帧内变换模式,对所述当前层中的节点进行正向变换,得到所述当前层中的节点的中间第二系数预测值;将所述中间第二系数预测值作为所述当前层中的节点的所述第二系数预测值。
在本申请的一些实施例中,在第四语法标识信息与第五语法标识信息均为第一值的情况下,所述第一语法标识信息为第一值;所述第四语法标识信息用于指示所述当前层的节点是否允许进行帧间预测;所述第五语法标识信息用于指示所述当前层的节点是否允许进行帧内预测;在第四语法标识信息与第五语法标识信息任意一个为第二值的情况下,所述第一语法标识信息为第二值。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为在所述第一语法标识信息指示所 述当前层不允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定第四语法标识信息;在所述第四语法标识信息指示所述当前层的节点不允许进行帧间预测的情况下,解析码流,确定第五语法标识信息;在所述第五语法标识信息指示所述当前层的节点允许进行帧内预测的情况下,根据区域自适应分层帧内变换模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为在所述第四语法标识信息指示所述当前层的节点允许进行帧间预测的情况下,根据区域自适应分层帧间变换模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为若所述第四语法标识信息的取值为第一值,则确定所述当前层的节点允许进行帧间预测;若所述第四语法标识信息的取值为第二值,则确定所述当前层的节点不允许进行帧间预测。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为若所述第五语法标识信息的取值为第一值,则确定所述当前层的节点允许进行帧间预测;若所述第五语法标识信息的取值为第二值,则确定所述当前层的节点不允许进行帧间预测。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为解析码流,确定第六语法标识信息的取值;所述第六语法标识信息用于指示所述当前层中的节点采用区域自适应分层帧间变换模式。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为确定所述当前层的相邻节点数目;其中,所述相邻节点包括邻域节点数目以及父节点邻域节点数目;在所述相邻节点数目大于或等于预设阈值的情况下,确定所述当前层的节点允许进行属性预测。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为基于所述当前层中各个节点的空间位置,确定各个所述节点的邻域节点;其中,所述节点的邻域节点至少包括:与所述节点共面的邻域节点和与所述节点共线的邻域节点。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为确定所述当前层中各个所述节点的父节点;基于各个所述节点的父节点的空间位置,确定各个所述节点的父节点邻域节点;其中,所述节点的父节点邻域节点至少包括:与所述节点的父节点共面的邻域节点和与所述节点的父节点共线的邻域节点。
在本申请的一些实施例中,所述第一确定部分1001,还被配置为对所述当前层中各个节点的邻域节点进行数量统计,确定所述当前层的邻域节点数目;对所述当前层中各个节点的父节点邻域节点进行数量统计,确定所述当前层的父节点邻域节点数目;将所述邻域节点数目和所述父节点邻域节点数目进行相加,得到所述当前层的所述相邻节点数目。
可以理解地,在本申请实施例中,“部分”可以是部分电路、部分处理器、部分程序或软件等等,当然也可以是模块,还可以是非模块化的。而且在本实施例中的各组成部分可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。
所述集成的单元如果以软件功能模块的形式实现并非作为独立的产品进行销售或使用时,可以存储在一个计算机可读取存储介质中,基于这样的理解,本实施例的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的全部或部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)或processor(处理器)执行本实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(Read Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、磁碟或者光盘等各种可以存储程序代码的介质。
因此,本申请实施例提供了一种计算机可读存储介质,应用于解码器1000,该计算机可读存储介质存储有计算机程序,所述计算机程序被第一处理器执行时实现前述实施例中任一项所述的方法。
基于上述解码器1000的组成以及计算机可读存储介质,参见图45,其示出了本申请实施例提供的解码器1000的具体硬件结构示意图。如图45所示,解码器1000可以包括:第一通信接口1101、第一存储器1102和第一处理器1103;各个组件通过第一总线系统1104耦合在一起。可理解,第一总线系统1104用于实现这些组件之间的连接通信。第一总线系统1104除包括数据总线之外,还包括电源总线、控制总线和状态信号总线。但是为了清楚说明起见,在图11中将各种总线都标为第一总线系统1104。其中,
第一通信接口1101,用于在与其他外部网元之间进行收发信息过程中,信号的接收和发送;
第一存储器1102,用于存储能够在第一处理器1103上运行的计算机程序;
第一处理器1103,用于在运行所述计算机程序时,执行:
在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;
在所述第一语法标识信息指示所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定所述当前层的目标解码模式;
根据所述目标解码模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值。
可以理解,本申请实施例中的第一存储器1102可以是易失性存储器或非易失性存储器,或可包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(Read-Only Memory,ROM)、可编程只读存储器(Programmable ROM,PROM)、可擦除可编程只读存储器(Erasable PROM,EPROM)、电可擦除可编程只读存储器(Electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(Random Access Memory,RAM),其用作外部高速缓存。通过示例性但不是限制性说明,许多形式的RAM可用,例如静态随机存取存储器(Static RAM,SRAM)、动态随机存取存储器(Dynamic RAM,DRAM)、同步动态随机存取存储器(Synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(Double Data Rate SDRAM,DDRSDRAM)、增强型同步动态随机存取存储器(Enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(Synchlink DRAM,SLDRAM)和直接内存总线随机存取存储器(Direct Rambus RAM,DRRAM)。本申请描述的系统和方法的第一存储器1102旨在包括但不限于这些和任意其它适合类型的存储器。
而第一处理器1103可能是一种集成电路芯片,具有信号的处理能力。在实现过程中,上述方法的各步骤可以通过第一处理器1103中的硬件的集成逻辑电路或者软件形式的指令完成。上述的第一处理器1103可以是通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现成可编程门阵列(Field Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件。可以实现或者执行本申请实施例中的公开的各方法、步骤及逻辑框图。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。结合本申请实施例所公开的方法的步骤可以直接体现为硬件译码处理器执行完成,或者用译码处理器中的硬件及软件模块组合执行完成。软件模块可以位于随机存储器,闪存、只读存储器,可编程只读存储器或者电可擦写可编程存储器、寄存器等本领域成熟的存储介质中。该存储介质位于第一存储器1102,第一处理器1103读取第一存储器1102中的信息,结合其硬件完成上述方法的步骤。
可以理解的是,本申请描述的这些实施例可以用硬件、软件、固件、中间件、微码或其组合来实现。对于硬件实现,处理单元可以实现在一个或多个专用集成电路(Application Specific Integrated Circuits,ASIC)、数字信号处理器(Digital Signal Processing,DSP)、数字信号处理设备(DSP Device,DSPD)、可编程逻辑设备(Programmable Logic Device,PLD)、现场可编程门阵列(Field-Programmable Gate Array,FPGA)、通用处理器、控制器、微控制器、微处理器、用于执行本申请所述功能的其它电子单元或其组合中。对于软件实现,可通过执行本申请所述功能的模块(例如过程、函数等)来实现本申请所述的技术。软件代码可存储在存储器中并通过处理器执行。存储器可以在处理器中或在处理器外部实现。
可选地,作为另一个实施例,第一处理器1103还配置为在运行所述计算机程序时,执行前述实施例中任一项所述的方法。
本实施例提供了一种解码器,在该解码器中,为每一层引入对应的属性解码模式,针对每一个层进行属性解码时,在解码端可以自适应地选择每一个层的目标解码模式,使得解码端使用所解析到的目标解码模式来对点云的属性进行属性重建,从而提升了点云属性的解码效率,进而提升了点云的解码性能。
在本申请的再一实施例中,基于前述实施例相同的发明构思,参见图46,其示出了本申请实施例提供的一种编码器的组成结构示意图。如图46所示,该编码器2000可以包括第二确定部分2001和编码部分2002;其中,
所述第二确定部分,用于在确定当前层的节点允许进行属性预测,以及所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,确定所述当前层的目标编码模式,并确定第一语法标识信息;所述第一语法标识信息用于指示所述当前层是否允许自适应选择帧间预测模式和/或帧内预测模式的情况;
所述编码部分,用于根据所述目标编码模式对所述当前层中的节点进行属性编码,确定所述当前层中的节点的属性重建值。
所述编码部分2002,被配置为若确定所述当前层允许自适应选择帧间预测模式和/或帧内预测模式,则确定所述第一语法标识信息的取值为第一值;若确定所述当前层不允许自适应选择帧间预测模式和/或帧内预测模式,则确定所述第一语法标识信息的取值为第二值。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为若确定当前系数组允许自适应选 择帧间预测模式和/或帧内预测模式,则确定所述第一语法标识信息的取值为第一值;所述当前系数组包括至少一个层,所述当前层为所述至少一个层的其中一层;若确定所述当前系数组不允许自适应选择帧间预测模式和/或帧内预测模式,则确定所述第一语法标识信息的取值为第二值。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为若确定所述当前层允许自适应选择帧间预测模式和/或帧内预测模式,则确定所述第一语法标识信息的取值为第一值;若确定所述当前层不允许自适应选择帧间预测模式和/或帧内预测模式,则确定所述第一语法标识信息的取值为第二值。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为根据所述目标编码模式,确定第二语法标识信息;将所述第二语法标识信息添加至属性块头信息参数集,并对所述属性块头信息参数集进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,所述目标编码模式包括区域自适应分层帧内变换模式、区域自适应分层帧间变换模式和区域自适应分层结合变换模式;其中,所述区域自适应分层帧内变换模式表征采用帧内预测模式对所述当前层的节点进行属性预测变换编码;所述区域自适应分层帧间变换模式表征采用帧间预测模式对所述当前层的节点进行属性预测变换编码;所述区域自适应分层结合变换模式表征采用帧内预测模式结合帧间预测模式对所述当前层的节点进行属性预测变换编码。
在本申请的一些实施例中,所述区域自适应分层帧间变换模式包括第一区域自适应分层帧间变换模式和第二区域自适应分层帧间变换模式;所述区域自适应分层结合变换模式包括第一区域自适应分层结合变换模式、第二区域自适应分层结合变换模式和第三区域自适应分层结合变换模式;其中,
所述第一区域自适应分层帧间变换模式表征利用节点的几何信息确定同位预测节点的方式对所述当前层的节点进行属性预测变换编码;
所述第二区域自适应分层结合变换模式表征利用参考帧的缓存确定同位预测节点的方式对所述当前层的节点进行属性预测变换编码;
所述第一区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和所述第一区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换编码;
所述第二区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和所述第二区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换编码;
所述第三区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式、所述第一区域自适应分层帧间变换模式和所述第二区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换编码。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为根据所述当前层采用帧间预测模式和/或帧内预测模式的所述目标解码模式,确定所述第二语法标识信息的取值。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为若所述当前层的所述目标编码模式为区域自适应分层帧内变换模式,则确定所述第二语法标识信息的取值为第三值;
若所述当前层的所述目标编码模式为第一区域自适应分层帧间变换模式,则确定所述第二语法标识信息的取值为第四值;
若所述当前层的所述目标编码模式为第二区域自适应分层帧间变换模式,则确定所述第二语法标识信息的取值为第五值;
若所述当前层的所述目标编码模式为第一区域自适应分层结合变换模式,则确定所述第二语法标识信息的取值为第六值;
若所述当前层的所述目标编码模式为第二区域自适应分层结合变换模式,则确定所述第二语法标识信息的取值为第七值;
若所述当前层的所述目标编码模式为第三区域自适应分层结合变换模式,则确定所述第二语法标识信息的取值为第八值。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为确定第三语法标识信息的取值;所述第三语法标识信息用于指示所述当前层所在的当前序列中所包含层的数量;获取所述当前层的索引值,若所述当前层的索引值为大于等于第九值且小于所述层的数量,则执行确定所述当前层的目标编码模式的步骤;若所述当前层的索引值为大于所述层的数量,则不执行确定所述当前层的目标编码模式的步骤;将所述第三语法标识信息进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为基于至少一种候选编码模式对所述当前层进行属性编码,确定所述至少一种候选编码模式各自的预编码结果;根据所述至少一种候选编码模式各自的预编码结果进行代价计算,确定所述至少一种候选编码模式各自的代价值;根据所述至少一种候选编码模式各自的代价值,从所述至少一种候选编码模式中确定所述当前层的所述目标编码模式。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为根据所述至少一种候选编码模式 各自的预编码结果进行率失真代价计算,确定所述至少一种候选编码模式各自的代价值。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为从所述至少一种候选编码模式各自的代价值中确定最小代价值;将所述最小代价值对应的候选编码模式确定为所述当前层的所述目标编码模式。
在本申请的一些实施例中,所述至少一种候选编码模式包括以下至少一种:区域自适应分层帧内变换模式、第一区域自适应分层帧间变换模式、第二区域自适应分层帧间变换模式、第一区域自适应分层结合变换模式、第二区域自适应分层结合变换模式、第三区域自适应分层结合变换模式。
在本申请的一些实施例中,所述编码部分2002,还被配置为确定所述当前层中的节点的属性预测值;根据所述目标编码模式对所述当前层中的节点的属性预测值进行正向变换,确定所述当前层中的节点的第一系数值和第二系数预测值;根据所述第二系数预测值确定所述当前层的节点对应的第二系数值;根据所述目标编码模式对所述当前层中的节点的所述第一系数值和所述第二系数值进行逆向变换,确定所述当前层中的节点的属性重建值。
在本申请的一些实施例中,所述编码部分2002,还被配置为确定所述当前层中的节点的相邻节点;其中,所述相邻节点包括邻域节点和父节点邻域节点;根据所述相邻节点对应的属性重建值以及所述当前层中的节点与所述相邻节点之间的几何距离进行线性拟合,确定所述当前层中的节点的属性预测值。
在本申请的一些实施例中,所述编码部分2002,还被配置为确定所述当前层中的节点的第二系数编码残差值;对所述第二系数编码残差值进行反量化处理,得到所述当前层中的节点的第二系数反量化残差值;根据所述当前层中的节点对应的第二系数预测值和所述第二系数反化残差值,确定所述当前层中的节点的第二系数值。
在本申请的一些实施例中,所述编码部分2002,还被配置为确定所述当前层中的节点的属性原始值;根据所述目标编码模式对所述当前层中的节点的所述属性原始值进行正向变换,确定所述当前层中的节点的所述第一系数值和第二系数原始值;根据所述当前层中的点的所述第二系数原始值和所述第二系数预测值,确定所述当前层中的节点的第二系数预测残差值;对所述第二系数残差值进行量化处理,得到所述当前层中的节点的第二系数量化残差值。
在本申请的一些实施例中,所述编码部分2002,还被配置为对所述当前层中的节点的所述第二系数量化残差值进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,在当前层的目标编码模式为所述区域自适应分层结合变换模式的情况下,所述编码部分2002,还被配置为根据所述区域自适应分层结合变换模式对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一中间预测值和第二中间预测值;将所述当前层中的节点的所述第一中间预测值和所述第二中间预测值进行相加,得到所述当前层中的节点的所述第二系数预测值。
在本申请的一些实施例中,所述编码部分2002,还被配置为在预设权重表中确定所述当前层对应的目标权重组合;其中,所述目标权重组合包括第一目标权重和第二目标权重;采用所述区域自适应分层帧内变换模式,对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一属性预测值;采用所述区域自适应分层帧间变换模式,对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第二属性预测值;将所述当前层中的节点的所述第一属性预测值与所述第一目标权重进行相乘,得到所述当前层中的节点的所述第一中间预测值;将所述当前层中的节点的所述第二属性预测值与所述第二目标权重进行相乘,得到所述当前层中的节点的所述第二中间预测值。
在本申请的一些实施例中,所述编码部分2002,还被配置为将所述目标权重组合在所述预设权重表中对应的权重索引值进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,所述编码部分2002,还被配置为确定所述当前层中的节点的属性预测值;根据所述目标编码模式对所述当前层中的节点的属性预测值进行正向变换,确定所述当前层中的节点的第一系数值和第二系数预测值;根据所述第二系数预测值确定所述当前层的节点对应的第二系数值;根据所述目标编码模式对所述当前层中的节点的所述第一系数值和所述第二系数值进行逆向变换,确定所述当前层中的节点的属性重建值。
在本申请的一些实施例中,所述编码部分2002,还被配置为确定所述当前层中的节点的相邻节点;其中,所述相邻节点包括邻域节点和父节点邻域节点;根据所述相邻节点对应的属性重建值以及所述当前层中的节点与所述相邻节点之间的几何距离进行线性拟合,确定所述当前层中的节点的属性预测值。
在本申请的一些实施例中,所述编码部分2002,还被配置为确定所述当前层中的节点的第二系数编码残差值;对所述第二系数编码残差值进行反量化处理,得到所述当前层中的节点的第二系数反量化残差值;根据所述当前层中的节点对应的第二系数预测值和所述第二系数反化残差值,确定所述当前层中的节点的第二系数值。
在本申请的一些实施例中,所述编码部分2002,还被配置为确定所述当前层中的节点的属性原始 值;根据所述目标编码模式对所述当前层中的节点的所述属性原始值进行正向变换,确定所述当前层中的节点的所述第一系数值和第二系数原始值;根据所述当前层中的点的所述第二系数原始值和所述第二系数预测值,确定所述当前层中的节点的第二系数预测残差值;对所述第二系数残差值进行量化处理,得到所述当前层中的节点的第二系数量化残差值。
在本申请的一些实施例中,所述编码部分2002,还被配置为对所述当前层中的节点的所述第二系数量化残差值进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,在当前层的目标编码模式为所述区域自适应分层结合变换模式的情况下,所述编码部分2002,还被配置为根据所述区域自适应分层结合变换模式对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一中间预测值和第二中间预测值;将所述当前层中的节点的所述第一中间预测值和所述第二中间预测值进行相加,得到所述当前层中的节点的所述第二系数预测值。
在本申请的一些实施例中,所述编码部分2002,还被配置为在预设权重表中确定所述当前层对应的目标权重组合;其中,所述目标权重组合包括第一目标权重和第二目标权重;采用所述区域自适应分层帧内变换模式,对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一属性预测值;采用所述区域自适应分层帧间变换模式,对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第二属性预测值;将所述当前层中的节点的所述第一属性预测值与所述第一目标权重进行相乘,得到所述当前层中的节点的所述第一中间预测值;将所述当前层中的节点的所述第二属性预测值与所述第二目标权重进行相乘,得到所述当前层中的节点的所述第二中间预测值。
在本申请的一些实施例中,所述编码部分2002,还被配置为将所述目标权重组合在所述预设权重表中对应的权重索引值进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,在当前层的目标编码模式为所述区域自适应分层帧间变换模式或所述区域自适应分层结合变换模式的情况下,所述编码部分2002,还被配置为在所述当前层中的节点的第二系数预测值为第十值的情况下,根据所述区域自适应分层帧内变换模式对所述当前层中的节点进行正向变换,得到所述当前层中的节点的中间第二系数预测值;将所述中间第二系数预测值作为所述当前层中的节点的所述第二系数预测值。
在本申请的一些实施例中,所述编码部分2002,还被配置为确定第四语法标识信息和第五语法标识信息的取值;所述第五语法标识信息用于指示所述当前层的节点是否允许进行帧内预测;在所述第四语法标识信息与所述第五语法标识信息均为第一值的情况下,则确定所述第一语法标识信息的取值为第一值;在第四语法标识信息与第五语法标识信息任意一个为第二值的情况下,则确定所述第一语法标识信息的取值为第二值;将所述第一语法标识信息进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为若确定所述当前层的节点允许进行帧间预测,则将第四语法标识信息的取值设置为第一值;若确定所述当前层的节点不允许进行帧间预测,则将所述第四语法标识信息的取值设置为第二值;将所述第四语法标识信息进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为若确定所述当前层的节点允许进行帧内预测,则将第五语法标识信息的取值设置为第一值;若确定所述当前层的节点不允许进行帧内预测,则将所述第五语法标识信息的取值设置为第二值;将所述第五语法标识信息进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为确定第六语法标识信息的取值;所述第六语法标识信息用于指示所述当前层中的节点采用区域自适应分层帧间变换模式;将所述第六语法标识信息进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为对所述当前层中的节点的所述目标编码模式进行编码处理,将所得到的编码比特写入码流。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为确定所述当前层的相邻节点数目;其中,所述相邻节点包括邻域节点数目以及父节点邻域节点数目;在所述相邻节点数目大于或等于预设阈值的情况下,确定所述当前层的节点允许进行属性预测。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为基于所述当前层中各个节点的空间位置,确定各个所述节点的邻域节点;其中,所述节点的邻域节点至少包括:与所述节点共面的邻域节点和与所述节点共线的邻域节点。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为确定所述当前层中各个所述节点的父节点;基于各个所述节点的父节点的空间位置,确定各个所述节点的父节点邻域节点;其中,所述节点的父节点邻域节点至少包括:与所述节点的父节点共面的邻域节点和与所述节点的父节点共线的邻域节点。
在本申请的一些实施例中,所述第二确定部分2001,还被配置为对所述当前层中各个节点的邻域节点进行数量统计,确定所述当前层的邻域节点数目;对所述当前层中各个节点的父节点邻域节点进行数量统计,确定所述当前层的父节点邻域节点数目;将所述邻域节点数目和所述父节点邻域节点数目进行相加,得到所述当前层的所述相邻节点数目。
可以理解地,在本实施例中,“部分”可以是部分电路、部分处理器、部分程序或软件等等,当然也可以是模块,还可以是非模块化的。而且在本实施例中的各组成部分可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。
所述集成的单元如果以软件功能模块的形式实现并非作为独立的产品进行销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本实施例提供了一种计算机可读存储介质,应用于编码器2000,该计算机可读存储介质存储有计算机程序,所述计算机程序被第二处理器执行时实现前述实施例中任一项所述的方法。
基于上述编码器2000的组成以及计算机可读存储介质,参见图47,其示出了本申请实施例提供的编码器2000的具体硬件结构示意图。如图47所示,编码器2000可以包括:第二通信接口2101、第二存储器2102和第二处理器2103;各个组件通过第二总线系统2104耦合在一起。可理解,第二总线系统2104用于实现这些组件之间的连接通信。第二总线系统2104除包括数据总线之外,还包括电源总线、控制总线和状态信号总线。但是为了清楚说明起见,在图21中将各种总线都标为第二总线系统2104。其中,
第二通信接口2101,用于在与其他外部网元之间进行收发信息过程中,信号的接收和发送;
第二存储器2102,用于存储能够在第二处理器2103上运行的计算机程序;
第二处理器2103,用于在运行所述计算机程序时,执行:
在确定当前层的节点允许进行属性预测的情况下,确定第一语法标识信息;
在确定当前层的节点允许进行属性预测,以及所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,确定所述当前层的目标编码模式,并确定第一语法标识信息;所述第一语法标识信息用于指示所述当前层是否允许自适应选择帧间预测模式和/或帧内预测模式的情况;
根据所述目标编码模式对所述当前层中的节点进行属性编码,确定所述当前层中的节点的属性重建值。
可选地,作为另一个实施例,第二处理器2103还配置为在运行所述计算机程序时,执行前述实施例中任一项所述的方法。
可以理解,第二存储器2102与第一存储器4602的硬件功能类似,第二处理器2103与第一处理器1103的硬件功能类似;这里不再详述。
本实施例提供了一种编码器,在该编码器中,为每一层引入对应的属性编码模式,针对每一个层进行属性编码时,在编码端可以自适应地选择每一个层的目标编码模式,使得编码端使用所解析到的目标编码模式来对点云的属性进行属性重建,从而提升了点云属性的编码效率,进而提升了点云的编码性能。
在本申请的再一实施例中,参见图48,其示出了本申请实施例提供的一种编解码系统的组成结构示意图。如图48所示,编解码系统3000可以包括解码器3001和编码器3002。
在本申请实施例中,解码器3001可以是前述实施例中任一项所述的解码器,编码器3002可以是前述实施例中任一项所述的编码器。
需要说明的是,在本申请中,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者装置不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者装置所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括该要素的过程、方法、物品或者装置中还存在另外的相同要素。
上述本申请实施例序号仅仅为了描述,不代表实施例的优劣。
本申请所提供的几个方法实施例中所揭露的方法,在不冲突的情况下可以任意组合,得到新的方法实施例。
本申请所提供的几个产品实施例中所揭露的特征,在不冲突的情况下可以任意组合,得到新的产品实施例。
本申请所提供的几个方法或设备实施例中所揭露的特征,在不冲突的情况下可以任意组合,得到新的方法实施例或设备实施例。
以上所述,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技 术领域的技术人员在本申请揭露的技术范围内,可轻易想到变化或替换,都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应以所述权利要求的保护范围为准。
工业实用性
本申请实施例中,在解码端,在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;在第一语法标识信息指示当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定当前层的目标解码模式;根据目标解码模式对当前层中的节点进行属性解码,确定当前层中的节点的属性重建值。在编码端,在确定当前层的节点允许进行属性预测,以及当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,确定当前层的目标编码模式,并确定第一语法标识信息;第一语法标识信息用于指示当前层是否允许自适应选择帧间预测模式和/或帧内预测模式的情况;根据目标编码模式对当前层中的节点进行属性编码,确定当前层中的节点的属性重建值。这样,通过为每一层引入对应的属性编码模式,针对每一个层进行属性编码时,在编码端可以自适应地选择每一个片的目标编码模式,并将目标编码模式传递给解码端,使得解码端使用所解析到的目标解码模式来对点云的属性进行属性重建,从而提升了点云属性的编解码效率,进而提升了点云的编解码性能。

Claims (63)

  1. 一种解码方法,应用于解码器,所述方法包括:
    在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;
    在所述第一语法标识信息指示所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定所述当前层的目标解码模式;
    根据所述目标解码模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值。
  2. 根据权利要求1所述的方法,其中,所述解析码流,确定第一语法标识信息,包括:
    若所述第一语法标识信息的取值为第一值,则确定当前系数组允许自适应选择帧间预测模式和/或帧内预测模式;所述当前系数组包括至少一个层,所述当前层为所述至少一个层的其中一层;
    若所述第一语法标识信息的取值为第二值,则确定所述当前系数组不允许自适应选择帧间预测模式和/或帧内预测模式。
  3. 根据权利要求1或2所述的方法,其中,所述解析码流,确定第一语法标识信息,包括:
    若所述第一语法标识信息的取值为第一值,则确定所述当前层允许自适应选择帧间预测模式和/或帧内预测模式;
    若所述第一语法标识信息的取值为第二值,则确定所述当前层不允许自适应选择帧间预测模式和/或帧内预测模式。
  4. 根据权利要求1至3任一项所述的方法,其中,所述解析码流,确定所述当前层的目标解码模式,包括:
    解码码流,确定属性块头信息参数集;
    从所述属性块头信息参数集中,确定第二语法标识信息;
    根据所述第二语法标识信息,确定所述当前层的所述目标解码模式。
  5. 根据权利要求4所述的方法,其中,所述目标解码模式包括区域自适应分层帧内变换模式、区域自适应分层帧间变换模式和区域自适应分层结合变换模式;其中,
    所述区域自适应分层帧内变换模式表征采用帧内预测模式对所述当前层的节点进行属性预测变换解码;所述区域自适应分层帧间变换模式表征采用帧间预测模式对所述当前层的节点进行属性预测变换解码;所述区域自适应分层结合变换模式表征采用帧内预测模式结合帧间预测模式对所述当前层的节点进行属性预测变换解码。
  6. 根据权利要求5所述的方法,其中,所述区域自适应分层帧间变换模式包括第一区域自适应分层帧间变换模式和第二区域自适应分层帧间变换模式;所述区域自适应分层结合变换模式包括第一区域自适应分层结合变换模式、第二区域自适应分层结合变换模式和第三区域自适应分层结合变换模式;其中,
    所述第一区域自适应分层帧间变换模式表征利用节点的几何信息确定同位预测节点的方式对所述当前层的节点进行属性预测变换解码;
    所述第二区域自适应分层帧间变换模式表征利用参考帧的缓存确定同位预测节点的方式对所述当前层的节点进行属性预测变换解码;
    所述第一区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和所述第一区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换解码;
    所述第二区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和所述第二区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换解码;
    所述第三区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式、所述第一区域自适应分层帧间变换模式和所述第二区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换解码。
  7. 根据权利要求4至6任一项所述的方法,其中,所述根据所述第二语法标识信息,确定所述当前层的所述目标解码模式,包括:
    根据所述第二语法标识信息的取值,确定所述当前层采用帧间预测模式和/或帧内预测模式的所述目标解码模式。
  8. 根据权利要求7所述的方法,其中,所述根据所述第二语法标识信息的取值,确定所述当前层采用帧间预测模式和/或帧内预测模式的所述目标解码模式,包括:
    若所述第二语法标识信息的取值为第三值,则确定所述当前层的所述目标解码模式为区域自适应分 层帧内变换模式;
    若所述第二语法标识信息的取值为第四值,则确定所述当前层的所述目标解码模式为第一区域自适应分层帧间变换模式;
    若所述第二语法标识信息的取值为第五值,则确定所述当前层的所述目标解码模式为第二区域自适应分层帧间变换模式;
    若所述第二语法标识信息的取值为第六值,则确定所述当前层的所述目标解码模式为第一区域自适应分层结合变换模式;
    若所述第二语法标识信息的取值为第七值,则确定所述当前层的所述目标解码模式为第二区域自适应分层结合变换模式;
    若所述第二语法标识信息的取值为第八值,则确定所述当前层的所述目标解码模式为第三区域自适应分层结合变换模式。
  9. 根据权利要求1所述的方法,其中,所述方法还包括:
    解析码流,确定第三语法标识信息的取值;所述第三语法标识信息用于指示所述当前层所在的当前序列中所包含的层的数量;
    获取所述当前层的索引值,若所述当前层的索引值为大于等于第九值且小于所述层的数量,则执行所述解析码流,确定所述当前层的目标解码模式的步骤;
    若所述当前层的索引值为大于所述层的数量,则不执行所述解析码流,确定所述当前层的目标解码模式的步骤。
  10. 根据权利要求1至9任一项所述的方法,其中,所述根据所述目标解码模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值,包括:
    确定所述当前层中的节点的属性预测值;
    根据所述目标解码模式对所述当前层中的节点的属性预测值进行正向变换,确定所述当前层中的节点的第一系数值和第二系数预测值;
    根据所述第二系数预测值确定所述当前层的节点对应的第二系数值;
    根据所述目标解码模式对所述当前层中的节点的所述第一系数值和所述第二系数值进行逆向变换,确定所述当前层中的节点的属性重建值。
  11. 根据权利要求10所述的方法,其中,所述确定所述当前层中的节点的属性预测值,包括:
    确定所述当前层中的节点的相邻节点;其中,所述相邻节点包括邻域节点和父节点邻域节点;
    根据所述相邻节点对应的属性重建值以及所述当前层中的节点与所述相邻节点之间的几何距离进行线性拟合,确定所述当前层中的节点的所述属性预测值。
  12. 根据权利要求10所述的方法,其中,所述根据所述第二系数预测值确定所述当前层的节点对应的第二系数值,包括:
    解码码流,确定所述当前层中的节点的第二系数解码残差值;
    对所述第二系数解码残差值进行反量化处理,得到所述当前层中的节点的第二系数反量化残差值;
    根据所述当前层中的节点对应的第二系数预测值和所述第二系数反化残差值,确定所述当前层中的节点的所述第二系数值。
  13. 根据权利要求10至12任一项所述的方法,其中,在当前层的目标解码模式为所述区域自适应分层结合变换模式的情况下,确定所述当前层中的节点的第二系数预测值,包括:
    根据所述区域自适应分层结合变换模式对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一中间预测值和第二中间预测值;
    将所述当前层中的节点的所述第一中间预测值和所述第二中间预测值进行相加,得到所述当前层中的节点的所述第二系数预测值。
  14. 根据权利要求13所述的方法,其中,所述根据所述区域自适应分层结合变换模式对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一中间预测值和第二中间预测值,包括:
    确定所述当前层的第一目标权重和第二目标权重;
    采用所述区域自适应分层帧内变换模式,对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一属性预测值;
    采用所述区域自适应分层帧间变换模式,对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第二属性预测值;
    将所述当前层中的节点的所述第一属性预测值与所述第一目标权重进行相乘,得到所述当前层中的节点的所述第一中间预测值;
    将所述当前层中的节点的所述第二属性预测值与所述第二目标权重进行相乘,得到所述当前层中的 节点的所述第二中间预测值。
  15. 根据权利要求14所述的方法,其中,所述确定所述当前层的第一目标权重和第二目标权重,包括:
    解析码流,确定权重索引值;
    在预设权重表中确定与所述权重索引对应的目标权重组合;其中,所述目标权重组合包括第一目标权重和第二目标权重。
  16. 根据权利要求10至15任一项所述的方法,其中,在当前层的目标解码模式为所述区域自适应分层帧间变换模式或所述区域自适应分层结合变换模式的情况下,在确定所述当前层中的节点的第二系数预测值之后,所述方法还包括:
    在所述当前层中的节点的第二系数预测值为第十值的情况下,采用所述区域自适应分层帧内变换模式,对所述当前层中的节点进行正向变换,得到所述当前层中的节点的中间第二系数预测值;
    将所述中间第二系数预测值作为所述当前层中的节点的所述第二系数预测值。
  17. 根据权利要求1所述的方法,其中,在第四语法标识信息与第五语法标识信息均为第一值的情况下,所述第一语法标识信息为第一值;所述第四语法标识信息用于指示所述当前层的节点是否允许进行帧间预测;所述第五语法标识信息用于指示所述当前层的节点是否允许进行帧内预测;
    在第四语法标识信息与第五语法标识信息任意一个为第二值的情况下,所述第一语法标识信息为第二值。
  18. 根据权利要求1至17任一项所述的方法,其中,所述方法还包括:
    在所述第一语法标识信息指示所述当前层不允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定第四语法标识信息;
    在所述第四语法标识信息指示所述当前层的节点不允许进行帧间预测的情况下,解析码流,确定第五语法标识信息;
    在所述第五语法标识信息指示所述当前层的节点允许进行帧内预测的情况下,根据区域自适应分层帧内变换模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值。
  19. 根据权利要求18所述的方法,其中,所述方法还包括:在所述第四语法标识信息指示所述当前层的节点允许进行帧间预测的情况下,根据区域自适应分层帧间变换模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值。
  20. 根据权利要求18或19所述的方法,其中,所述方法还包括:
    若所述第四语法标识信息的取值为第一值,则确定所述当前层的节点允许进行帧间预测;
    若所述第四语法标识信息的取值为第二值,则确定所述当前层的节点不允许进行帧间预测。
  21. 根据权利要求18或19所述的方法,其中,所述方法还包括:
    若所述第五语法标识信息的取值为第一值,则确定所述当前层的节点允许进行帧间预测;
    若所述第五语法标识信息的取值为第二值,则确定所述当前层的节点不允许进行帧间预测。
  22. 根据权利要求1所述的方法,其中,所述方法还包括:
    解析码流,确定第六语法标识信息的取值;所述第六语法标识信息用于指示所述当前层中的节点采用区域自适应分层帧间变换模式。
  23. 根据权利要求1至16任一项所述的方法,其中,所述方法还包括:
    确定所述当前层的相邻节点数目;其中,所述相邻节点包括邻域节点数目以及父节点邻域节点数目;
    在所述相邻节点数目大于或等于预设阈值的情况下,确定所述当前层的节点允许进行属性预测。
  24. 根据权利要求23所述的方法,其中,所述方法还包括:
    基于所述当前层中各个节点的空间位置,确定各个所述节点的邻域节点;
    其中,所述节点的邻域节点至少包括:与所述节点共面的邻域节点和与所述节点共线的邻域节点。
  25. 根据权利要求23所述的方法,其中,所述方法还包括:
    确定所述当前层中各个所述节点的父节点;
    基于各个所述节点的父节点的空间位置,确定各个所述节点的父节点邻域节点;其中,所述节点的父节点邻域节点至少包括:与所述节点的父节点共面的邻域节点和与所述节点的父节点共线的邻域节点。
  26. 根据权利要求23至25任一项所述的方法,其中,所述确定所述当前层的相邻节点数目,包括:
    对所述当前层中各个节点的邻域节点进行数量统计,确定所述当前层的邻域节点数目;
    对所述当前层中各个节点的父节点邻域节点进行数量统计,确定所述当前层的父节点邻域节点数目;
    将所述邻域节点数目和所述父节点邻域节点数目进行相加,得到所述当前层的所述相邻节点数目。
  27. 一种编码方法,应用于编码器,所述方法包括:
    在确定当前层的节点允许进行属性预测,以及所述当前层允许自适应选择帧间预测模式和/或帧内 预测模式的情况下,确定所述当前层的目标编码模式,并确定第一语法标识信息;所述第一语法标识信息用于指示所述当前层是否允许自适应选择帧间预测模式和/或帧内预测模式的情况;
    根据所述目标编码模式对所述当前层中的节点进行属性编码,确定所述当前层中的节点的属性重建值。
  28. 根据权利要求27所述的方法,其中,所述确定第一语法标识信息,包括:
    若确定当前系数组允许自适应选择帧间预测模式和/或帧内预测模式,则确定所述第一语法标识信息的取值为第一值;所述当前系数组包括至少一个层,所述当前层为所述至少一个层的其中一层;
    若确定所述当前系数组不允许自适应选择帧间预测模式和/或帧内预测模式,则确定所述第一语法标识信息的取值为第二值。
  29. 根据权利要求27或28所述的方法,其中,所述确定第一语法标识信息,包括:
    若确定所述当前层允许自适应选择帧间预测模式和/或帧内预测模式,则确定所述第一语法标识信息的取值为第一值;
    若确定所述当前层不允许自适应选择帧间预测模式和/或帧内预测模式,则确定所述第一语法标识信息的取值为第二值。
  30. 根据权利要求27至29任一项所述的方法,其中,所述方法还包括:
    根据所述目标编码模式,确定第二语法标识信息;
    将所述第二语法标识信息添加至属性块头信息参数集,并对所述属性块头信息参数集进行编码处理,将所得到的编码比特写入码流。
  31. 根据权利要求27所述的方法,其中,所述目标编码模式包括区域自适应分层帧内变换模式、区域自适应分层帧间变换模式和区域自适应分层结合变换模式;其中,
    所述区域自适应分层帧内变换模式表征采用帧内预测模式对所述当前层的节点进行属性预测变换编码;所述区域自适应分层帧间变换模式表征采用帧间预测模式对所述当前层的节点进行属性预测变换编码;所述区域自适应分层结合变换模式表征采用帧内预测模式结合帧间预测模式对所述当前层的节点进行属性预测变换编码。
  32. 根据权利要求31所述的方法,其中,所述区域自适应分层帧间变换模式包括第一区域自适应分层帧间变换模式和第二区域自适应分层帧间变换模式;所述区域自适应分层结合变换模式包括第一区域自适应分层结合变换模式、第二区域自适应分层结合变换模式和第三区域自适应分层结合变换模式;其中,
    所述第一区域自适应分层帧间变换模式表征利用节点的几何信息确定同位预测节点的方式对所述当前层的节点进行属性预测变换编码;
    所述第二区域自适应分层帧间变换模式表征利用参考帧的缓存确定同位预测节点的方式对所述当前层的节点进行属性预测变换编码;
    所述第一区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和所述第一区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换编码;
    所述第二区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式和所述第二区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换编码;
    所述第三区域自适应分层结合变换模式表征采用结合区域自适应分层帧内变换模式、所述第一区域自适应分层帧间变换模式和所述第二区域自适应分层帧间变换模式对所述当前层的节点进行属性预测变换编码。
  33. 根据权利要求30至32任一项所述的方法,其中,所述根据所述目标编码模式,确定第二语法标识信息,包括:
    根据所述当前层采用帧间预测模式和/或帧内预测模式的所述目标解码模式,确定所述第二语法标识信息的取值。
  34. 根据权利要求33所述的方法,其中,所述根据所述当前层采用帧间预测模式和/或帧内预测模式的所述目标解码模式,确定所述第二语法标识信息的取值,包括:
    若所述当前层的所述目标编码模式为区域自适应分层帧内变换模式,则确定所述第二语法标识信息的取值为第三值;
    若所述当前层的所述目标编码模式为第一区域自适应分层帧间变换模式,则确定所述第二语法标识信息的取值为第四值;
    若所述当前层的所述目标编码模式为第二区域自适应分层帧间变换模式,则确定所述第二语法标识信息的取值为第五值;
    若所述当前层的所述目标编码模式为第一区域自适应分层结合变换模式,则确定所述第二语法标识 信息的取值为第六值;
    若所述当前层的所述目标编码模式为第二区域自适应分层结合变换模式,则确定所述第二语法标识信息的取值为第七值;
    若所述当前层的所述目标编码模式为第三区域自适应分层结合变换模式,则确定所述第二语法标识信息的取值为第八值。
  35. 根据权利要求27所述的方法,其中,所述方法还包括:
    确定第三语法标识信息的取值;所述第三语法标识信息用于指示所述当前层所在的当前序列中所包含层的数量;
    获取所述当前层的索引值,若所述当前层的索引值为大于等于第九值且小于所述层的数量,则执行确定所述当前层的目标编码模式的步骤;
    若所述当前层的索引值为大于所述层的数量,则不执行确定所述当前层的目标编码模式的步骤;
    将所述第三语法标识信息进行编码处理,将所得到的编码比特写入码流。
  36. 根据权利要求27至35任一项所述的方法,其中,所述确定所述当前层的目标编码模式,包括:
    基于至少一种候选编码模式对所述当前层进行属性编码,确定所述至少一种候选编码模式各自的预编码结果;
    根据所述至少一种候选编码模式各自的预编码结果进行代价计算,确定所述至少一种候选编码模式各自的代价值;
    根据所述至少一种候选编码模式各自的代价值,从所述至少一种候选编码模式中确定所述当前层的所述目标编码模式。
  37. 根据权利要求36所述的方法,其中,所述根据所述至少一种候选编码模式各自的预编码结果进行代价计算,确定所述至少一种候选编码模式各自的代价值,包括:
    根据所述至少一种候选编码模式各自的预编码结果进行率失真代价计算,确定所述至少一种候选编码模式各自的代价值。
  38. 根据权利要求36所述的方法,其中,所述根据所述至少一种候选编码模式各自的代价值,从所述至少一种候选编码模式中确定所述当前层的所述目标编码模式,包括:
    从所述至少一种候选编码模式各自的代价值中确定最小代价值;
    将所述最小代价值对应的候选编码模式确定为所述当前层的所述目标编码模式。
  39. 根据权利要求36至38任一项所述的方法,其中,所述至少一种候选编码模式包括以下至少一种:区域自适应分层帧内变换模式、第一区域自适应分层帧间变换模式、第二区域自适应分层帧间变换模式、第一区域自适应分层结合变换模式、第二区域自适应分层结合变换模式、第三区域自适应分层结合变换模式。
  40. 根据权利要求27至39任一项所述的方法,其中,所述根据所述目标编码模式对所述当前层中的节点进行属性编码,确定所述当前层中的节点的属性重建值,包括:
    确定所述当前层中的节点的属性预测值;
    根据所述目标编码模式对所述当前层中的节点的属性预测值进行正向变换,确定所述当前层中的节点的第一系数值和第二系数预测值;
    根据所述第二系数预测值确定所述当前层的节点对应的第二系数值;
    根据所述目标编码模式对所述当前层中的节点的所述第一系数值和所述第二系数值进行逆向变换,确定所述当前层中的节点的所述属性重建值。
  41. 根据权利要求40所述的方法,其中,所述确定所述当前层中的节点的属性预测值,包括:
    确定所述当前层中的节点的相邻节点;其中,所述相邻节点包括邻域节点和父节点邻域节点;
    根据所述相邻节点对应的属性重建值以及所述当前层中的节点与所述相邻节点之间的几何距离进行线性拟合,确定所述当前层中的节点的属性预测值。
  42. 根据权利要求40所述的方法,其中,所述根据所述第二系数预测值确定所述当前层的节点对应的第二系数值,包括:
    确定所述当前层中的节点的第二系数编码残差值;
    对所述第二系数编码残差值进行反量化处理,得到所述当前层中的节点的第二系数反量化残差值;
    根据所述当前层中的节点对应的第二系数预测值和所述第二系数反化残差值,确定所述当前层中的节点的第二系数值。
  43. 根据权利要求42所述的方法,其中,所述确定所述当前层中的节点的第二系数编码残差值,包括:
    确定所述当前层中的节点的属性原始值;
    根据所述目标编码模式对所述当前层中的节点的所述属性原始值进行正向变换,确定所述当前层中的节点的所述第一系数值和第二系数原始值;
    根据所述当前层中的点的所述第二系数原始值和所述第二系数预测值,确定所述当前层中的节点的第二系数预测残差值;
    对所述第二系数残差值进行量化处理,得到所述当前层中的节点的所述第二系数量化残差值。
  44. 根据权利要求43所述的方法,其中,所述方法还包括:
    对所述当前层中的节点的所述第二系数量化残差值进行编码处理,将所得到的编码比特写入码流。
  45. 根据权利要求40至44任一项所述的方法,其中,在当前层的目标编码模式为所述区域自适应分层结合变换模式的情况下,确定所述当前层中的节点的第二系数预测值,包括:
    根据所述区域自适应分层结合变换模式对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一中间预测值和第二中间预测值;
    将所述当前层中的节点的所述第一中间预测值和所述第二中间预测值进行相加,得到所述当前层中的节点的所述第二系数预测值。
  46. 根据权利要求45所述的方法,其中,所述根据所述区域自适应分层结合变换模式对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一中间预测值和第二中间预测值,包括:
    在预设权重表中确定所述当前层对应的目标权重组合;其中,所述目标权重组合包括第一目标权重和第二目标权重;
    采用所述区域自适应分层帧内变换模式,对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第一属性预测值;
    采用所述区域自适应分层帧间变换模式,对所述当前层中的节点进行正向变换,确定所述当前层中的节点的第二属性预测值;
    将所述当前层中的节点的所述第一属性预测值与所述第一目标权重进行相乘,得到所述当前层中的节点的所述第一中间预测值;
    将所述当前层中的节点的所述第二属性预测值与所述第二目标权重进行相乘,得到所述当前层中的节点的所述第二中间预测值。
  47. 根据权利要求46所述的方法,其中,将所述目标权重组合在所述预设权重表中对应的权重索引值进行编码处理,将所得到的编码比特写入码流。
  48. 根据权利要求40至47任一项所述的方法,其中,在当前层的目标编码模式为所述区域自适应分层帧间变换模式或所述区域自适应分层结合变换模式的情况下,在确定所述当前层中的节点的第二系数预测值之后,所述方法还包括:
    在所述当前层中的节点的第二系数预测值为第十值的情况下,根据所述区域自适应分层帧内变换模式对所述当前层中的节点进行正向变换,得到所述当前层中的节点的中间第二系数预测值;
    将所述中间第二系数预测值作为所述当前层中的节点的所述第二系数预测值。
  49. 根据权利要求27所述的方法,其中,所述方法还包括:
    确定第四语法标识信息和第五语法标识信息的取值;所述第五语法标识信息用于指示所述当前层的节点是否允许进行帧内预测;
    在所述第四语法标识信息与所述第五语法标识信息均为第一值的情况下,则确定所述第一语法标识信息的取值为第一值;
    在第四语法标识信息与第五语法标识信息任意一个为第二值的情况下,则确定所述第一语法标识信息的取值为第二值;
    将所述第一语法标识信息进行编码处理,将所得到的编码比特写入码流。
  50. 根据权利要求27至49任一项所述的方法,其中,所述方法还包括:
    若确定所述当前层的节点允许进行帧间预测,则将第四语法标识信息的取值设置为第一值;
    若确定所述当前层的节点不允许进行帧间预测,则将所述第四语法标识信息的取值设置为第二值;
    将所述第四语法标识信息进行编码处理,将所得到的编码比特写入码流。
  51. 根据权利要求27至49任一项所述的方法,其中,所述方法还包括:
    若确定所述当前层的节点允许进行帧内预测,则将第五语法标识信息的取值设置为第一值;
    若确定所述当前层的节点不允许进行帧内预测,则将所述第五语法标识信息的取值设置为第二值;
    将所述第五语法标识信息进行编码处理,将所得到的编码比特写入码流。
  52. 根据权利要求27至49任一项所述的方法,其中,所述方法还包括:
    确定第六语法标识信息的取值;所述第六语法标识信息用于指示所述当前层中的节点采用区域自适应分层帧间变换模式;
    将所述第六语法标识信息进行编码处理,将所得到的编码比特写入码流。
  53. 根据权利要求27至49任一项所述的方法,其中,所述方法还包括:
    对所述当前层中的节点的所述目标编码模式进行编码处理,将所得到的编码比特写入码流。
  54. 根据权利要求27至49任一项所述的方法,其中,所述方法还包括:
    确定所述当前层的相邻节点数目;其中,所述相邻节点包括邻域节点数目以及父节点邻域节点数目;
    在所述相邻节点数目大于或等于预设阈值的情况下,确定所述当前层的节点允许进行属性预测。
  55. 根据权利要求54所述的方法,其中,所述方法还包括:
    基于所述当前层中各个节点的空间位置,确定各个所述节点的邻域节点;
    其中,所述节点的邻域节点至少包括:与所述节点共面的邻域节点和与所述节点共线的邻域节点。
  56. 根据权利要求54所述的方法,其中,所述方法还包括:
    确定所述当前层中各个所述节点的父节点;
    基于各个所述节点的父节点的空间位置,确定各个所述节点的父节点邻域节点;其中,所述节点的父节点邻域节点至少包括:与所述节点的父节点共面的邻域节点和与所述节点的父节点共线的邻域节点。
  57. 根据权利要求54至56任一项所述的方法,其中,所述确定所述当前层的相邻节点数目,包括:
    对所述当前层中各个节点的邻域节点进行数量统计,确定所述当前层的邻域节点数目;
    对所述当前层中各个节点的父节点邻域节点进行数量统计,确定所述当前层的父节点邻域节点数目;
    将所述邻域节点数目和所述父节点邻域节点数目进行相加,得到所述当前层的所述相邻节点数目。
  58. 一种码流,其中,所述码流是根据待编码信息进行比特编码生成的;其中,待编码信息包括下述至少一项:第一语法标识信息的取值、第二语法标识信息的取值、第三语法标识信息的取值、第四语法标识信息的取值、第五语法标识信息的取值、所述当前层中的节点对应的权重索引值、所述当前层中的节点的第二系数量化残差值;
    其中,所述第一语法标识信息用于指示所述当前层是否允许自适应选择帧间预测模式和/或帧内预测模式,所述第二语法标识信息用于指示当前层的目标编码模式,所述第三语法标识信息的取值用于指示所述当前层所在的当前序列中所包含层的数量,所述第四语法标识信息的取值用于指示所述当前层的节点是否允许进行帧间预测,所述第五语法标识信息用于指示所述当前层的节点是否允许进行帧内预测,所述第六语法标识信息用于指示所述当前层中的节点采用区域自适应分层帧间变换模式,所述权重索引值用于指示所述当前层中的节点对应的目标权重组合在预设权重表中对应的索引值。
  59. 一种解码器,所述解码器包括第一确定部分和解码部分;其中,
    所述第一确定部分,用于在确定当前层的节点允许进行属性预测的情况下,解析码流,确定第一语法标识信息;在所述第一语法标识信息指示所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,解析码流,确定所述当前层的目标解码模式;
    所述解码部分,用于根据所述目标解码模式对所述当前层中的节点进行属性解码,确定所述当前层中的节点的属性重建值。
  60. 一种解码器,所述解码器包括第一存储器和第一处理器;其中,
    所述第一存储器,用于存储能够在所述第一处理器上运行的计算机程序;
    所述第一处理器,用于在运行所述计算机程序时,执行如权利要求1至26中任一项所述的方法。
  61. 一种编码器,所述编码器包括第二确定部分和编码部分;其中,
    所述第二确定部分,用于在确定当前层的节点允许进行属性预测,以及所述当前层允许自适应选择帧间预测模式和/或帧内预测模式的情况下,确定所述当前层的目标编码模式,并确定第一语法标识信息;所述第一语法标识信息用于指示所述当前层是否允许自适应选择帧间预测模式和/或帧内预测模式的情况;
    所述编码部分,用于根据所述目标编码模式对所述当前层中的节点进行属性编码,确定所述当前层中的节点的属性重建值。
  62. 一种编码器,所述编码器包括第二存储器和第二处理器;其中,
    所述第二存储器,用于存储能够在所述第二处理器上运行的计算机程序;
    所述第二处理器,用于在运行所述计算机程序时,执行如权利要求27至57中任一项所述的方法。
  63. 一种计算机可读存储介质,其中,所述计算机可读存储介质存储有计算机程序,所述计算机程序被执行时实现如权利要求1至26中任一项所述的方法、或者实现如权利要求27至57中任一项所述的方法。
PCT/CN2023/106200 2023-07-06 2023-07-06 编解码方法、码流、编码器、解码器以及存储介质 Ceased WO2025007360A1 (zh)

Priority Applications (3)

Application Number Priority Date Filing Date Title
CN202380099308.3A CN121359460A (zh) 2023-07-06 2023-07-06 编解码方法、码流、编码器、解码器以及存储介质
PCT/CN2023/106200 WO2025007360A1 (zh) 2023-07-06 2023-07-06 编解码方法、码流、编码器、解码器以及存储介质
US19/423,905 US20260113436A1 (en) 2023-07-06 2025-12-17 Coding method, decoding method, bit stream, coder, decoder, and storage medium

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2023/106200 WO2025007360A1 (zh) 2023-07-06 2023-07-06 编解码方法、码流、编码器、解码器以及存储介质

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US19/423,905 Continuation US20260113436A1 (en) 2023-07-06 2025-12-17 Coding method, decoding method, bit stream, coder, decoder, and storage medium

Publications (1)

Publication Number Publication Date
WO2025007360A1 true WO2025007360A1 (zh) 2025-01-09

Family

ID=94171018

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/106200 Ceased WO2025007360A1 (zh) 2023-07-06 2023-07-06 编解码方法、码流、编码器、解码器以及存储介质

Country Status (3)

Country Link
US (1) US20260113436A1 (zh)
CN (1) CN121359460A (zh)
WO (1) WO2025007360A1 (zh)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170347100A1 (en) * 2016-05-28 2017-11-30 Microsoft Technology Licensing, Llc Region-adaptive hierarchical transform and entropy coding for point cloud compression, and corresponding decompression
WO2022145214A1 (ja) * 2020-12-28 2022-07-07 ソニーグループ株式会社 情報処理装置および方法
WO2023287265A1 (ko) * 2021-07-16 2023-01-19 엘지전자 주식회사 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법
CN116233388A (zh) * 2021-12-03 2023-06-06 维沃移动通信有限公司 点云编、解码处理方法、装置、编码设备及解码设备

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170347100A1 (en) * 2016-05-28 2017-11-30 Microsoft Technology Licensing, Llc Region-adaptive hierarchical transform and entropy coding for point cloud compression, and corresponding decompression
WO2022145214A1 (ja) * 2020-12-28 2022-07-07 ソニーグループ株式会社 情報処理装置および方法
WO2023287265A1 (ko) * 2021-07-16 2023-01-19 엘지전자 주식회사 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법
CN116233388A (zh) * 2021-12-03 2023-06-06 维沃移动通信有限公司 点云编、解码处理方法、装置、编码设备及解码设备

Also Published As

Publication number Publication date
US20260113436A1 (en) 2026-04-23
CN121359460A (zh) 2026-01-16

Similar Documents

Publication Publication Date Title
WO2025007360A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
US20260113485A1 (en) Encoding method, decoding method, code stream, encoder, decoder, and storage medium
US20260046395A1 (en) Encoding and decoding methods, encoder, decoder, bitstream, and storage medium
US20260046450A1 (en) Encoding and decoding method, code stream, encoder, decoder and storage medium
WO2025076672A1 (zh) 编解码方法、编码器、解码器、码流以及存储介质
WO2025010600A9 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2024216476A1 (zh) 编解码方法、编码器、解码器、码流以及存储介质
US20260129188A1 (en) Method for encoding, method for decoding, and storage medium
WO2025010604A1 (zh) 点云编解码方法、编码器、解码器、码流以及存储介质
WO2024234132A9 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2025076668A9 (zh) 编解码方法、编码器、解码器以及存储介质
WO2025010601A9 (zh) 编解码方法、编码器、解码器、码流以及存储介质
WO2025007349A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2025145433A1 (zh) 点云编解码方法、编解码器、码流以及存储介质
WO2024207481A1 (zh) 编解码方法、编码器、解码器、码流以及存储介质
WO2025076663A1 (zh) 编解码方法、编解码器以及存储介质
WO2024207456A1 (zh) 编解码方法、编码器、解码器、码流以及存储介质
WO2025145330A1 (zh) 点云编解码方法、编解码器、码流以及存储介质
WO2024212038A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2026007149A1 (zh) 点云编解码方法、码流、编码器、解码器以及存储介质
WO2024212043A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2025147915A1 (zh) 点云编解码方法、编解码器、码流以及存储介质
WO2024212042A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2024212045A1 (zh) 编解码方法、码流、编码器、解码器以及存储介质
WO2024148598A1 (zh) 编解码方法、编码器、解码器以及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23944087

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE