WO2024256314A1 - Signaling for implicit neural representation reconstruction - Google Patents
Signaling for implicit neural representation reconstruction Download PDFInfo
- Publication number
- WO2024256314A1 WO2024256314A1 PCT/EP2024/065877 EP2024065877W WO2024256314A1 WO 2024256314 A1 WO2024256314 A1 WO 2024256314A1 EP 2024065877 W EP2024065877 W EP 2024065877W WO 2024256314 A1 WO2024256314 A1 WO 2024256314A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- representation network
- implicit
- neural representation
- inr
- indication
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/70—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
Definitions
- the present embodiments generally relate to image, video and/or 3D scene compression using Implicit Neural Representation (INR).
- INR Implicit Neural Representation
- the present embodiments relate to a method and an apparatus for encoding, decoding, transmitting metadata used by a decoder for reconstructing a signal encoded using an INR network.
- BACKGROUND To achieve high compression efficiency, image and video coding schemes usually employ prediction and transform to leverage spatial and temporal redundancy in the video content. Emerging technology makes use of neural networks.
- Implicit Neural Representation aims at parameterizing a function which takes coordinates as inputs and outputs values of a signal at these coordinates.
- INR can be used for instance for compressing image, videos or 3D objects or scene. It can also apply to any type of signal.
- Approaches are known to construct an INR network for encoding 2D or 3D images.
- Approaches are also known to compress a neural network.
- any image/video decoder could not reconstruct the image or video for display based only on the INR. Additional information is needed to address any image/video decoder for reconstructing an output signal.
- SUMMARY a method for signaling one or more syntax elements for using an INR decoder is provided.
- the one or more syntax elements provides for reconstructing at least one part of a signal representative of a scene using at least one Implicit Neural Representation network representative of the at least one part of the signal representative of the scene.
- the method also comprises encoding the INR network.
- an apparatus for signaling one or more syntax elements for using an INR decoder comprises one or more processors operable to signal the one or more syntax elements that provides for reconstructing at least one part of a signal representative of a scene using at least one Implicit Neural Representation network representative of the at least one part of a signal representative of a scene.
- the apparatus is also operable to encode the INR network.
- a method for decoding one or more syntax elements for using an INR decoder is provided.
- the one or more syntax elements provides for reconstructing at least one part of a signal representative of a scene using at least one Implicit Neural Representation network representative of the at least one part of a signal representative of a scene.
- the method also comprises decoding the INR network.
- the method also comprises reconstructing the at least one part of a signal representative of a scene using the one or more syntax elements and the Implicit Neural Representation network.
- an apparatus for decoding one or more syntax elements for using an INR decoder is provided.
- the apparatus comprises one or more processors operable to decode the one or more syntax elements that provides for reconstructing at least one part of a signal representative of a scene using at least one Implicit Neural Representation network representative of the at least one part of a signal representative of a scene.
- the apparatus is also operable to decode the INR network.
- the apparatus is also operable to reconstruct the at least one part of a signal representative of a scene using the one or more syntax elements and the Implicit Neural Representation network. Further embodiments that can be used alone or in combination are described herein.
- One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform the method for signaling/decoding one or more syntax elements for using an INR decoder according to any of the embodiments described herein.
- One or more of the present embodiments also provide a non-transitory computer readable medium and/or a computer readable storage medium having stored thereon instructions for signaling/decoding one or more syntax elements for using an INR decoder according to the methods described herein.
- One or more embodiments also provide a computer readable storage medium having stored thereon a bitstream generated according to the methods described herein.
- One or more embodiments also provide a method and apparatus for transmitting or receiving the bitstream generated according to the methods described above.
- FIG.1 illustrates an example of a neural network for Implicit Neural Representation.
- FIG. 2 illustrates an example of a method for encoding information used by a decoder to reconstruct at least one part of a signal representative of a scene from an encoded INR, according to an embodiment.
- FIG. 3 illustrates an example of a method for decoding information used by a decoder to reconstruct at least one part of a signal representative of a scene from an encoded INR, according to an embodiment.
- FIG. 4 illustrates an example of a method for reconstructing at least one part of a signal representative of a scene from an encoded INR, according to an embodiment.
- FIG.5 illustrates an example of a method for generating coordinates to be used by a decoder to reconstruct at least one part of a signal representative of a scene from an encoded INR, according to an embodiment.
- FIG. 6 illustrates an example of a method for performing an inference of the INR network, according to an embodiment.
- FIG.7 illustrates an example of a method for reconstructing the at least one part of a signal representative of a scene from an output of the INR network, according to an embodiment.
- FIG. 8 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented.
- FIG. 9 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented, according to another embodiment.
- FIG. 10 shows two remote devices communicating over a communication network in accordance with an example of the present principles.
- FIG.11 shows the syntax of a signal in accordance with an example of the present principles.
- DETAILED DESCRIPTION This application describes a variety of aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that may sound limiting. However, this is for purposes of clarity in description, and does not limit the application or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well. The aspects described and contemplated in this application can be implemented in many different forms. FIGs.
- FIGs. 1-11 below provide some embodiments, but other embodiments are contemplated and the discussion of FIGs. 1-11 does not limit the breadth of the implementations.
- the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, the terms “image,” “picture” and “frame” may be used interchangeably.
- Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and/or use of specific steps and/or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc.
- first decoding and a “second decoding”.
- first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
- aspects described in this application can be used individually or in combination. At least one of the aspects generally relates to image, video or 3D data encoding and decoding using Implicit Neural Representation.
- At least one of the aspects described herein relates to using an Implicit Neural Representation for encoding/decoding any signal representative of a scene.
- At least one other aspect generally relates to transmitting a bitstream generated or encoded.
- These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium having stored thereon instructions for encoding or decoding data signal representative of a scene according to any of the methods described, and/or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described. While not standardized yet, MPEG group standardization is exploring new Neural Compression technologies.
- FIG.1 illustrates an example of a neural network used for Implicit Neural Representation (INR).
- INR parameterizes a signal as a function 100, which takes coordinates 110 as input and outputs values 120 of a signal at these coordinates.
- the inputs 110 can be pixel coordinates ( ⁇ , ⁇ ) and the INR may output 120 the color values ( ⁇ , ⁇ , ⁇ ) or ( ⁇ , ⁇ , ⁇ ) of the input pixel.
- the input coordinates may be modified by a transformation before being used as input for the neural network. This transformation can be a Fourier mapping, coordinate transformation, normalization etc.
- the INR can be used to reconstruct a signal by computing the signal values for every necessary coordinate inputs. It can be used to upsample a signal by generating output for input coordinates corresponding to the upsampled pixels, for example the mean of the coordinates between two consecutive pixels for upsampling by a factor of 2.
- An INR network 100 is typically a neural network, composed of multiple neural layers, such as fully connected layers. In FIG. 1, the network has four layers. Intermediate outputs are represented by circles. Each neural layer can be described as a function that first multiplies the input by a tensor, adds a vector called the bias and then applies a nonlinear function on the resulting values. The shape (and other characteristics) of the tensor and the type of non- linear functions are called the architecture of the network.
- weights The values of the tensor and the bias are denoted herein by the term “weights”.
- the architecture and the parameters define a “model”.
- the notation ⁇ ⁇ is used to denote an INR function parameterized by ⁇ .
- a typical process to encode a signal using an INR is as follows. First the weights ⁇ (or a subset of them) of the INR network are optimized to reconstruct the signal. Next, these weights are optionally encoded to create the output bitstream.
- ⁇ could be any differentiable distortion measure, such as mean squared error as in the second equation.
- M and N are the width and height of an image. Other metrics such as LPIPS (learned perceptual image patch similarity) can also be used in this case.
- the optimization of the weights ⁇ is typically performed by a machine learning approach such as a batch gradient descent method.
- ⁇ ⁇ is evaluated at all relevant coordinates. These coordinates can be selected at decoding.
- a typical choice would be all pixel coordinates for an image or video. As an example, for a 256x256 pixel image, these coordinates could be all pairs ( ⁇ , ⁇ ) for all ⁇ ⁇ ⁇ 0,1, ... ,255 ⁇ and ⁇ ⁇ ⁇ 0,1, ... ,255 ⁇ .
- Other choices are possible, for example to upsample, downsample or extend the original image.
- the scientific literature describes many approaches to construct an INR network to encode a 2D or 3D image.
- NNC MPEG Neural Network Compression
- FIG.2 illustrates an example of a method 200 for encoding information used by a decoder to reconstruct at least one part of a signal representative of a scene from an encoded INR, according to an embodiment.
- a bitstream is constructed that contains an INR network, optimized and encoded by off-the-shelves approaches and an encoding of information necessary for a decoder to use the INR network to reconstruct the signal.
- the INR is encoded using for instance a NNC encoder.
- the INR is representative of at least one part of a signal representative of a scene.
- FIG.3 illustrates an example of a method 300 for decoding information used by a decoder to reconstruct at least one part of a signal representative of a scene from an encoded INR, according to an embodiment.
- bitstream encoded by the method described in relation with FIG.2 is decoded.
- one or more syntax elements SE are decoded from the bitstream.
- the syntax elements are provided to a decoder for reconstructing the at least one part of a signal representative of a scene using the syntax elements and the Implicit Neural Representation network.
- Some embodiments are provided below that describe examples of syntax elements mentioned above.
- some embodiments are provided below for reconstructing the at least one part of a signal representative of a scene using the decoded syntax elements.
- the bitstream should contain signaling for some or all the following elements.
- a type of the INR used for representing the at least one part of a signal representative of a scene is signaled.
- This signaling is related to the INR used, for example a single network for a whole image or scene, one network per channel or per groups of channels (such as one network for Y and one for UV), one network per patch in an image or per sub- volume of a scene etc.
- information relating to the input of the INR network is signaled. These elements describe information necessary to construct the input of the INR.
- the dimensions of the domain can be signaled, for example an image is a 2D signal. This describes the dimensions of the output image or scene.
- Input coordinate normalization can be signaled. Coordinates normalization refers to the range of the coordinates expected by the INR.
- one INR may expect input values in the range [0,1] 2 , in the range (1, ... , ⁇ ⁇ ⁇ ⁇ h) ⁇ (1, ... , h ⁇ ⁇ ⁇ h ⁇ ) or in a range [ ⁇ 1,1] 2 .
- Input coordinate transformation(s) can be signaled. Input coordinates are often transformed by a function before being fed to the neural network. In order to reconstruct an image, it is necessary to signal the transformation(s) that must be used. If multiple transformations are present, their order may also be signaled. Maximum upsampling size may be signaled. This information can be optional. It may be interesting to signal the maximum recommended upsampling possible with this INR.
- the maximum upsampling could be defined in various ways, for example by the largest upsampling rate that does not increase distortion by a set amount or the upsampling rate where INR upsampling is better than an upsampling algorithm available in the decoder.
- information relating to models are signaled.
- the models are usually integerized. That is the weights are quantized and expressed as fixed point values, and the floating point operations are transformed into their integer counterpart.
- activation layers such as sin(x)
- the range and bitdepth of the inputs/outputs need to be precisely described.
- information relating to the signal reconstruction are signaled.
- the output format can be signaled.
- This signaling indicates the format of an image encoding generated by the INR network, for example RGB or YUV. Additional signaling may be necessary depending on the encoding.
- the output range of the INR network can be signaled, for example [0,255] or [0,1].
- Table 1 describes some of the syntax elements that might be needed to describe the input and output of the network.
- o inr_model_type 0: the model takes as input a vector of dimension d (specified by - inr_input_dimension_count_minus1) and output an output. Input is sampled uniformly in the range specified for each dimension.
- o inr_model_type 1: Input is sampled using an 32-times downsampling function in the range specified for each dimension, at the center of each 32x32 patch. Alternatively, an additional parameter may specify the downsampling range.
- o inr_model_type 2: the INR contains two neural networks, the first one is applied before the coordinate transforms and the second one is applied afterwards.
- o Xxxx other approaches could also be specified.
- - inr_inference_type tag describing the type of computation: integer or floating point. This may impact the syntax of some elements. As an example, a value of 0 means integer and 1 floating point computation.
- - inr_input_dimension_count_minus1 integer describing the number of dimensions minus 1 of the input. For example, for an image INR model, the input has 2 dimensions (width and height), thus inr_input_dimension_count_minus1 is 1.
- the flag inr_input_zero_centered can be replaced by an explicit encoding of the offset of the zero value. For example, a model aimed at encoding an image of size WxH, with a sub-pixel accuracy of 1/8 of pixel (i.e.
- inr_input_transformation_type[i] tag describing the type of transformations applied to the input.
- the type 0 is let to define custom type.
- o inr_input_transformation_type[i] 1: hyperspherical coordinates transform
- o inr_input_transformation_type[i] 2: Fourier mapping using a custom Fourier mapping matrix. This type value necessitates that inr_inference_type is float.
- custom matrix is described as follows: ⁇ inr_Fourier_mapping_coefficient_count_minus1[i]: integer describing the number of Fourier mapping coefficient transformations minus 1. ⁇ inr_Fourier_mapping_coefficient[i][j]: value of a Fourier mapping coefficient.
- o inr_output_type 1: RGB output with each component encoded on 8 bits in the range [0,255]
- o inr_output_type 2: RGB output with each component encoded on 10 bits in the range [0,1023]
- o inr_output_type 3: YUV output with each component encoded on 8 bits in the range [0,255],
- o inr_output_type 4: YUV output with each component encoded on 10 bits in the range [0,1023],
- o inr_output_type 5: RGBA output with each component encoded on 8 bits in the range [0,255], where A represents an alpha value, o etc.
- LUT_table_s initialize()
- LUT_table_s[x][i] sin_int_rep(range,Fourier_mapping_coef[i])
- LUT_table_c[x][i] cos_int_rep(range,Fourier_mapping_coef[i]) ⁇
- len(Fourier_mapping_coef) returns the length of the input set Fourier_mapping_coef
- get_range(x) returns the range of real values mapped to the integer representation
- sin_int_rep(range,coef) is a function that returns an integer representation corresponding to the value of the sinus function on this interval.
- LUT_table_c[x][i]: LUT_table_s[LUT_complementary_angle table [x]][i].
- the syntax elements mentioned with FIG.2 and 3 can also comprise network configuration information.
- Neural network encoding can be signaled.
- This element signals how the INR neural network is encoded in the bitstream. It may for example be a value associated to a specific format such as NNC or Open Neural Network Exchange format and/or a particular version of a format.
- Inference engine configuration can be signaled. It may be interesting to add signaling related to the configuration of the inference engine, which could for example include the precision to use for the operations, the memory necessary to store the INR network or the inference engine to use.
- a value equals to 0 indicates that this SEI message contains an ISO/IEC 15938-17 bitstream
- a value of 1 indicates that this SEI message contains a neural network encoded by a format identified by the tag URI inr_tag_uri.
- inr_tag_uri contains a tag URI with syntax and semantics as specified in IETF RFC 4151 identifying the format and associated information about the INR network or an update encoded here.
- inr_uri contains a URI with syntax and semantics as specified in IETF Internet Standard 66 identifying the neural network used as an INR network or an update relative to the network.
- inr_complexity_info_present_flag specifies whether syntax element indicating the complexity of the INR network are present or not.
- a value equal to 1 specifies that one or more syntax elements that indicate the complexity of the INR network are present
- inr_complexity_info_present_flag 0 specifies that no syntax element indicating the complexity of the INR network is present.
- inr_parameter_type_idc 0 indicates that the neural network uses only integer parameters.
- inr_parameter_type_flag 1 indicates that the neural network may use floating point or integer parameters.
- inr_parameter_type_idc equal to 2 indicates that the neural network uses only binary parameters.
- inr_parameter_type_idc equal to 3 is reserved for future use.
- inr_log2_parameter_bit_length_minus3 0
- 1, 2, and 3 indicates that the neural network does not use parameters of bit length greater than 8, 16, 32, and 64, respectively.
- inr_parameter_type_idc is present and inr_log2_parameter_bit_length_minus3 is not present the neural network does not use parameters of bit length greater than 1.
- inr_num_parameters_idc indicates the maximum number of neural network parameters for the INR network in units of a power of 2 (or another power of 2).
- inr_num_parameters_idc 0 indicates that the maximum number of neural network parameters is unknown.
- the value inr_num_parameters_idc shall be in the range of 0 to 63, inclusive.
- maxNumParameters ( 2 ⁇ inr_num_parameters_idc ) ⁇ 1
- the number of neural network parameters of the post-processing filter shall be less than or equal to maxNumParameters.
- inr_num_mac_operations_idc greater than 0 indicates that the maximum number of multiply-accumulate operations per inference of INR network is less than or equal to inr_num_mac_operations_idc.
- inr_num_mac_operations_idc equal to 0 indicates that the maximum number of multiply-accumulate operations of the network is unknown.
- inr_num_mac_operations_idc shall be in the range of 0 to 2 32 ⁇ 1, inclusive.
- inr_total_kilobyte_size greater than 0 indicates a total size in kilobytes required to store the uncompressed parameters for the neural network. The total size in bits is a number equal to or greater than the sum of bits used to store each parameter.
- inr_total_kilobyte_size is the total size in bits divided by 8000, rounded up.
- inr_total_kilobyte_size equal to 0 indicates that the total size required to store the parameters for the neural network is unknown.
- the value of inr_total_kilobyte_size shall be in the range of 0 to 2 32 ⁇ 1, inclusive.
- inr_reserved_zero_bit_b shall be equal to 0 in bitstreams conforming to the syntax provided herein. Decoders shall ignore INR SEI messages in which inr_reserved_zero_bit_b is not equal to 0.
- inr_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO/IEC 15938-17.
- the byte sequence inr_payload_byte[i] for all present values of i shall be a complete bitstream that conforms to ISO/IEC 15938-17.
- the syntax elements provided above are only examples, other syntax elements can be added or some can be removed.
- FIG.4 illustrates a high level overview of an example method 400 that a decoder may rely on to decode a bitstream 410 that comprises an encoded INR representative of at least one part of a signal representative of a scene.
- Embodiments are described here for a signal representative of an image. But the described embodiments apply to any other kinds of signal representative of a scene.
- the bitstream also comprises one or more syntax elements according to an embodiment as described above in relation with Table 1 or 2 for example.
- the input bitstream 410 or parts of this bitstream is/are fed to the modules of the process.
- the module 420 extracts information from the bitstream, generates and outputs the set of input coordinates for the INR networks.
- the module 430 performs the INR inference. It is responsible for decoding the INR network and for generating the output of the INR network or in other words to associate the output of the INR network to each set of input coordinates.
- the module 440 combines these outputs to reconstruct the image. This may involve transformations on these outputs to match the requested output format. But this module is also responsible to properly arrange these values into an image format.
- FIG.5 illustrates one possible embodiment of the input generation module 520. In a step 510, the size of the original image signaled in the bitstream is taken as input 515 and used to generate input coordinates corresponding to the original image.
- these input coordinates are typically pairs ( ⁇ , ⁇ ) positions where ⁇ and ⁇ respectively range from one to the width or height of the image.
- the input is typically the coordinates of the block.
- Step 530 takes as input the signals in the bitstream related to the coordinate transformation(s) 535 and applies these transformations to the coordinates provided at 510.
- Step 550 arranges the coordinates based on the type of INR 555 signaled in the bitstream.
- FIG.6 illustrates one possible embodiment of the INR inference module 430.
- the INR network encoding signaled in the bitstream 615 is used as input to prepare the decoding.
- This piece of information is used to configure the INR decoder module 620.
- This module 620 takes as input the bitstream of the INR network 625 (that is the encoded INR) and outputs the decoded INR network.
- the module 630 uses the inference engine configuration 635 and the decoded INR network to initialize the inference engine 640.
- This module 630 may involve steps such as reserving computational resources, loading the decoded model into the memory of the processor that will perform inference, transforming the network into the format expected by the inference engine, adapting the weights of the network for the inference engine etc. It may also involve defining which coordinates are used in which networks if multiple networks are used in the INR. In that case, the input 635 of the module 630 may also include the signals for the type of INR used.
- This initialized inference engine 640 takes as input the coordinates 560. These coordinates are fed to the INR network and inference is performed to compute values associated to these pixels. Several variations are possible for the inference.
- inference can be performed on one set of coordinates at a time, on multiple sets of coordinates at the same time or on multiple processors in parallel.
- This module outputs the coordinates and the associated pixel values 660 that have been computed.
- FIG.7 illustrates one possible embodiment of the inference output combination module 440.
- a module 710 receives the coordinates and associated pixels values 660. It may also use as input (716) the type of INR used and the output format of the networks. This module is responsible for reordering the pixel values in a traditional image format, for example in a row- major or column-major ordered matrix. If necessary, this module may also take as input the type of INR used.
- a module 720 takes as input the range of the outputs of the network 726. It then performs a coordinate change on the pixel values to obtain values lying in a traditional range for images, for examples [0,255] or [0,1].
- the final output 730 is the decoded picture. This output may be further modified to obtain a different image format, for example from YUV to RGB or the opposite. The process described above is only one example.
- FIG. 8 illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented.
- System 800 may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application.
- Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers.
- Elements of system 800 may be embodied in a single integrated circuit, multiple ICs, and/or discrete components.
- the processing and encoder/decoder elements of system 800 are distributed across multiple ICs and/or discrete components.
- the system 800 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and/or output ports.
- the system 800 is configured to implement one or more of the aspects described in this application.
- the system 800 includes at least one processor 810 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application.
- Processor 810 may include embedded memory, input output interface, and various other circuitries as known in the art.
- the system 800 includes at least one memory 820 (e.g., a volatile memory device, and/or a non-volatile memory device).
- System 800 includes a storage device 840, which may include non-volatile memory and/or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and/or optical disk drive.
- the storage device 840 may include an internal storage device, an attached storage device, and/or a network accessible storage device, as non-limiting examples.
- System 800 includes an encoder/decoder module 830 configured, for example, to process data to provide an INR representative of at least one part of a signal representative of a scene and/or an encoded INR representative of at least one part of a signal representative of a scene or a decoded INR representative of at least one part of a signal representative of a scene and/or at least one part of the signal reconstructed from the decoded INR, and the encoder/decoder module 830 may include its own processor and memory.
- the encoder/decoder module 830 represents module(s) that may be included in a device to perform the encoding and/or decoding functions. As is known, a device may include one or both of the encoding and decoding modules.
- encoder/decoder module 830 may be implemented as a separate element of system 800 or may be incorporated within processor 810 as a combination of hardware and software as known to those skilled in the art.
- Program code to be loaded onto processor 810 or encoder/decoder 830 to perform the various aspects described in this application may be stored in storage device 840 and subsequently loaded onto memory 820 for execution by processor 810.
- one or more of processor 810, memory 820, storage device 840, and encoder/decoder module 830 may store one or more of various items during the performance of the processes described in this application.
- Such stored items may include, but are not limited to, the input image, video, 3D data, weights of the INR, the decoded image, decoded video, decoded 3D data or portions of the decoded data, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
- memory inside of the processor 810 and/or the encoder/decoder module 830 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding.
- a memory external to the processing device for example, the processing device may be either the processor 810 or the encoder/decoder module 830) is used for one or more of these functions.
- the external memory may be the memory 820 and/or the storage device 840, for example, a dynamic volatile memory and/or a non-volatile flash memory.
- an external non-volatile flash memory is used to store the operating system of a television
- the input to the elements of system 800 may be provided through various input devices as indicated in block 805.
- Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and/or (iv) a High Definition Multimedia Interface (HDMI) input terminal.
- RF radio frequency
- COMP Component
- USB Universal Serial Bus
- HDMI High Definition Multimedia Interface
- the input devices of block 805 have associated respective input processing elements as known in the art.
- the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets.
- the RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers.
- the RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband.
- the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band.
- Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter.
- the RF portion includes an antenna.
- the USB and/or HDMI terminals may include respective interface processors for connecting system 800 to other electronic devices across USB and/or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 810 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 810 as necessary.
- the demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 810, and encoder/decoder 830 operating in combination with the memory and storage elements to process the data stream as necessary for presentation on an output device.
- Various elements of system 800 may be provided within an integrated housing, Within the integrated housing, the various elements may be interconnected and transmit data therebetween using suitable connection arrangement 815, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
- the system 800 includes communication interface 850 that enables communication with other devices via communication channel 890.
- the communication interface 850 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 890.
- the communication interface 850 may include, but is not limited to, a modem or network card and the communication channel 890 may be implemented, for example, within a wired and/or a wireless medium.
- Data is streamed to the system 800, in various embodiments, using a Wi-Fi network such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers).
- IEEE 802.11 IEEE refers to the Institute of Electrical and Electronics Engineers.
- the Wi-Fi signal of these embodiments is received over the communications channel 890 and the communications interface 850 which are adapted for Wi-Fi communications.
- the communications channel 890 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications.
- inventions provide streamed data to the system 800 using a set-top box that delivers the data over the HDMI connection of the input block 805. Still other embodiments provide streamed data to the system 800 using the RF connection of the input block 805. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network.
- the system 800 may provide an output signal to various output devices, including a display 865, speakers 875, and other peripheral devices 885.
- the display 865 of various embodiments includes one or more of, for example, a touchscreen display, an organic light- emitting diode (OLED) display, a curved display, and/or a foldable display.
- OLED organic light- emitting diode
- the display 865 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other device.
- the display 865 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop).
- the other peripheral devices 885 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and/or a lighting system.
- DVR digital versatile disc
- Various embodiments use one or more peripheral devices 885 that provide a function based on the output of the system 800. For example, a disk player performs the function of playing the output of the system 800.
- control signals are communicated between the system 800 and the display 865, speakers 875, or other peripheral devices 885 using signaling such as AV.Link, CEC, or other communications protocols that enable device-to-device control with or without user intervention.
- the output devices may be communicatively coupled to system 800 via dedicated connections through respective interfaces 860, 870, and 880. Alternatively, the output devices may be connected to system 800 using the communications channel 890 via the communications interface 850.
- the display 865 and speakers 875 may be integrated in a single unit with the other components of system 800 in an electronic device, for example, a television.
- the display interface 860 includes a display driver, for example, a timing controller (T Con) chip.
- the display 865 and speaker 875 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 805 is part of a separate set-top box.
- the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
- the embodiments can be carried out by computer software implemented by the processor 810 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits.
- the memory 820 can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples.
- the processor 810 can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
- FIG. 9 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented, according to another embodiment.
- FIG. 9 shows one embodiment of an apparatus 900 for encoding or decoding metadata used by an INR decoder according to any one of the embodiments described herein.
- the apparatus comprises Processor 910 and can be interconnected to a memory 920 through at least one port. Both Processor 910 and memory 920 can also have one or more additional interconnections to external connections.
- Processor 910 is also configured to encode at least one Implicit Neural Representation network representative of at least one part of a signal representative of a scene, and signal one or more syntax elements providing for using the encoded Implicit Neural Representation network to reconstruct the at least one part of a signal representative of a scene, using any one of the embodiments described herein.
- the processor 910 is configured to decode the one or more syntax elements providing for using the Implicit Neural Representation network representative of at least one part of a signal representative of a scene and reconstruct the at least one part of a signal representative of a scene using the one or more syntax elements and the Implicit Neural Representation network, using any one of the embodiments described herein.
- the processor 910 is configured using a computer program product comprising code instructions that implements any one of embodiments described herein. In an embodiment, illustrated in FIG.
- the device A comprises a processor in relation with memory RAM and ROM which are configured to implement a method for encoding one or more syntax elements for using an INR decoder, as described with FIG.1-7 and the device B comprises a processor in relation with memory RAM and ROM which are configured to implement a method for decoding one or more syntax elements for using an INR decoder as described in relation with FIG 1-7.
- the network is a broadcast network, adapted to broadcast/transmit a coded INR and one or more syntax elements from device A to decoding devices including the device B.
- the coded INR and the one or more syntax elements are transmitted in a same signal.
- the coded INR and the one or more syntax elements are transmitted separately in distinct signals.
- FIG. 11 shows an example of the syntax of a signal transmitted over a packet-based transmission protocol.
- Each transmitted packet P comprises a header H and a payload PAYLOAD.
- the payload PAYLOAD may comprise one or more syntax elements for using an INR decoder, according to any one of the embodiments described above.
- the one or more syntax elements comprise at least one of an indication relating to the Implicit Neural Representation network, an information providing for constructing an input of the Implicit Neural Representation network, or an information relating to an output of the Implicit Neural Representation network.
- the indication relating to the Implicit Neural Representation network comprises at least one of a model type of the Implicit Neural Representation network, a type of computation of the Implicit Neural Representation network, a format of an encoding of the Implicit Neural Representation network, an indication of a tag URI identifying a format of the Implicit Neural Representation network, an indication of a URI identifying the Implicit Neural Representation network, an indicator indicating whether complexity information relating to the Implicit Neural Representation network is present or not, an indication of a type of parameters used by the Implicit Neural Representation network, an indication relating to a bit length of parameters used by the Implicit Neural Representation network, an indicator indicating a maximum number of parameters for the Implicit Neural Representation network, an indication of a maximum number of multiply-accumulate operations per inference of the Implicit Neural Representation network, or an indication of a size for storing uncompressed parameters of the Implicit Neural Representation
- the information providing for constructing an input of the Implicit Neural Representation network comprises at least one of an indication of a number of dimension of the input, an indication of a range of a dimension of the input, an indication of a quantizer of a dimension of the input, an indication of an offset of a zero value in the range, an indication of a number of transformation applied to the input, an indication of a type of transformation applied to the input, or an indication of one or more parameters of the transformation.
- Various implementations involve decoding. “Decoding”, as used in this application, can encompass all or part of the processes performed, for example, on a received encoded INR in order to produce a final output suitable for display.
- such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding.
- processes also, or alternatively, include processes performed by a decoder of various implementations described in this application, for example, entropy decoding a sequence of binary symbols to reconstruct image, video or 3D data.
- syntax elements as used herein are descriptive terms. As such, they do not preclude the use of other syntax element names. This disclosure has described various pieces of information, such as for example syntax, that can be transmitted or stored, for example.
- This information can be packaged or arranged in a variety of manners, including for example manners common in image, video or neural network standards such as putting the information into an SPS, a PPS, a NAL unit, a header (for example, a NAL unit header, or a slice header), or an SEI message.
- Other manners are also available, including for example manners common for system level or application level standards such as putting the information into one or more of the following: a. SDP (session description protocol), a format for describing multimedia communication sessions for the purposes of session announcement and session invitation, for example as described in RFCs and used in conjunction with RTP (Real-time Transport Protocol) transmission.
- SDP session description protocol
- RTP Real-time Transport Protocol
- DASH MPD Media Presentation Description
- a Descriptor is associated to a Representation or collection of Representations to provide additional characteristic to the content Representation.
- RTP header extensions for example as used during RTP streaming.
- ISO Base Media File Format for example as used in OMAF and using boxes which are object-oriented building blocks defined by a unique type identifier and length also known as 'atoms' in some specifications.
- HLS HTTP live Streaming
- a manifest can be associated, for example, to a version or collection of versions of a content to provide characteristics of the version or collection of versions.
- FIG. 1 When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method/process.
- the implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program).
- An apparatus can be implemented in, for example, appropriate hardware, software, and firmware.
- a processor which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device.
- Processors also include communication devices, such as, for example, computers, cell phones, portable/personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
- PDAs portable/personal digital assistants
- this application may refer to “determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory. Further, this application may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
- this application may refer to “receiving” various pieces of information.
- Receiving is, as with “accessing”, intended to be a broad term.
- Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory).
- “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
- any of the following “and/or”, and “at least one of”, for example, in the cases of “A/B”, “A and/or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B).
- such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C).
- This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
- the word “signal” refers to, among other things, indicating something to a corresponding decoder.
- the same parameter is used at both the encoder side and the decoder side.
- an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter.
- signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways.
- one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
- implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted.
- the information can include, for example, instructions for performing a method, or data produced by one of the described implementations.
- a signal can be formatted to carry the bitstream of a described embodiment.
- Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal.
- the formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream.
- the information that the signal carries can be, for example, analog or digital information.
- the signal can be transmitted over a variety of different wired or wireless links, as is known.
- the signal can be stored on a processor- readable medium.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
Description
Claims
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24731365.3A EP4728746A1 (en) | 2023-06-13 | 2024-06-10 | Signaling for implicit neural representation reconstruction |
| CN202480039292.1A CN121359461A (en) | 2023-06-13 | 2024-06-10 | Signal notifications used for implicit neural representation remodeling |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP23305937.7 | 2023-06-13 | ||
| EP23305937 | 2023-06-13 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024256314A1 true WO2024256314A1 (en) | 2024-12-19 |
Family
ID=87060528
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2024/065877 Ceased WO2024256314A1 (en) | 2023-06-13 | 2024-06-10 | Signaling for implicit neural representation reconstruction |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4728746A1 (en) |
| CN (1) | CN121359461A (en) |
| WO (1) | WO2024256314A1 (en) |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2024074373A1 (en) * | 2022-10-04 | 2024-04-11 | Interdigital Ce Patent Holdings, Sas | Quantization of weights in a neural network based compression scheme |
-
2024
- 2024-06-10 WO PCT/EP2024/065877 patent/WO2024256314A1/en not_active Ceased
- 2024-06-10 EP EP24731365.3A patent/EP4728746A1/en active Pending
- 2024-06-10 CN CN202480039292.1A patent/CN121359461A/en active Pending
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2024074373A1 (en) * | 2022-10-04 | 2024-04-11 | Interdigital Ce Patent Holdings, Sas | Quantization of weights in a neural network based compression scheme |
Non-Patent Citations (5)
| Title |
|---|
| BANG GUN ET AL: "Implicit neural visual representation compression of 3D scenes", SPIE, 1000 20TH ST. BELLINGHAM WA 98225-6705 USA, vol. 12592, 25 March 2023 (2023-03-25), pages 125921A - 125921A, XP060174969, DOI: 10.1117/12.2669420 * |
| EMILIEN DUPONT ET AL: "COIN: COmpression with Implicit Neural representations", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 10 April 2021 (2021-04-10), XP081929782 * |
| HAO LI ET AL: "[INVR] EE1.2: Exploration experiments of 2D INVR methods in inter configuration", no. m63133, 17 April 2023 (2023-04-17), XP030310170, Retrieved from the Internet <URL:https://dms.mpeg.expert/doc_end_user/documents/142_Antalya/wg11/m63133-v1-%5BINVR%5DExplorationexperimentsof2DINVRmethodsininterconfiguration.zip [INVR] Exploration experiments of 2D INVR methods in inter configuration.docx> [retrieved on 20230417] * |
| KIRCHHOFFER HEINER ET AL: "Overview of the Neural Network Compression and Representation (NNR) Standard", IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 1 January 2021 (2021-01-01), USA, pages 1 - 1, XP055831747, ISSN: 1051-8215, Retrieved from the Internet <URL:https://ieeexplore.ieee.org/ielx7/76/4358651/09478787.pdf?tp=&arnumber=9478787&isnumber=4358651&ref=aHR0cHM6Ly9pZWVleHBsb3JlLmllZWUub3JnL2Fic3RyYWN0L2RvY3VtZW50Lzk0Nzg3ODc=> DOI: 10.1109/TCSVT.2021.3095970 * |
| YANNICK STR\"UMPLER ET AL: "Implicit Neural Representations for Image Compression", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 8 December 2021 (2021-12-08), XP091115347 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121359461A (en) | 2026-01-16 |
| EP4728746A1 (en) | 2026-04-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP4168940A1 (en) | Systems and methods for encoding/decoding a deep neural network | |
| US20250139835A1 (en) | A method and an apparatus for encoding/decoding a 3d mesh | |
| EP4677841A1 (en) | Coding unit based implicit neural representation (inr) | |
| WO2023046463A1 (en) | Methods and apparatuses for encoding/decoding a video | |
| US20250254331A1 (en) | Video encoding and decoding using operations constraint | |
| EP4728746A1 (en) | Signaling for implicit neural representation reconstruction | |
| EP4672755A1 (en) | SIGNALING FOR FEAST-BASED IMPLICIT NEURAL REPRESENTATION RECONSTRUCTION | |
| US20260122260A1 (en) | Carriage of multiple parameter sets in a media file | |
| EP4664881A1 (en) | Efficient compression of coding tree unit based implicit neural representation with neural network coding standard | |
| US20250142118A1 (en) | A method and an apparatus for encoding/decoding attributes of a 3d object | |
| WO2025168360A1 (en) | Multiscale dictionary learning and training of inr network | |
| US20260122277A1 (en) | Methods to describe the high-level syntax design of a bitstream carrying data coded using learning-based codec for point cloud content | |
| WO2025011935A1 (en) | Approximating implicit neural representation through learnt dictionary atoms | |
| WO2024052134A1 (en) | Methods and apparatuses for encoding and decoding a point cloud | |
| EP4676058A1 (en) | Encoding partition-based inr (implicit neural representation) | |
| US20260122280A1 (en) | A coding method or apparatus signaling an indication of camera parameters | |
| WO2025168361A1 (en) | Updated dictionary-driven implicit neural representation for image and video compression | |
| WO2024189204A1 (en) | Methods and apparatuses for encoding and decoding a point cloud | |
| WO2026096499A1 (en) | Carriage of multiple parameter sets in a media file | |
| WO2025140843A1 (en) | Multiple frequency fourier mapping for implicit neural representation based compression | |
| WO2025155854A1 (en) | Carriage of coded base mesh and displacement data of video-based dynamic mesh coding in isobmff media containers | |
| WO2024163481A1 (en) | A method and an apparatus for encoding/decoding at least one part of an image using multi-level context model | |
| WO2023222521A1 (en) | Sei adapted for multiple conformance points | |
| WO2025153454A1 (en) | Signaling supplementary information related to attributes in v3c bitstream and basemesh bitstream | |
| WO2026008513A1 (en) | Video specific dictionary learning for implicit neural compression |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24731365 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202517121123 Country of ref document: IN |
|
| WWP | Wipo information: published in national office |
Ref document number: 202517121123 Country of ref document: IN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024731365 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2024731365 Country of ref document: EP Effective date: 20260113 |
|
| ENP | Entry into the national phase |
Ref document number: 2024731365 Country of ref document: EP Effective date: 20260113 |
|
| ENP | Entry into the national phase |
Ref document number: 2024731365 Country of ref document: EP Effective date: 20260113 |
|
| ENP | Entry into the national phase |
Ref document number: 2024731365 Country of ref document: EP Effective date: 20260113 |
|
| WWP | Wipo information: published in national office |
Ref document number: 2024731365 Country of ref document: EP |



