EP4602807A1 - Apparatus and method for integrating neural-network post-filter supplemental enhancement information with iso base media file format - Google Patents

Apparatus and method for integrating neural-network post-filter supplemental enhancement information with iso base media file format

Info

Publication number
EP4602807A1
EP4602807A1 EP23793066.4A EP23793066A EP4602807A1 EP 4602807 A1 EP4602807 A1 EP 4602807A1 EP 23793066 A EP23793066 A EP 23793066A EP 4602807 A1 EP4602807 A1 EP 4602807A1
Authority
EP
European Patent Office
Prior art keywords
track
sei
sample group
sample
nnpfc
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23793066.4A
Other languages
German (de)
French (fr)
Inventor
Miska Matias Hannuksela
Kashyap KAMMACHI SREEDHAR
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Nokia Technologies Oy
Original Assignee
Nokia Technologies Oy
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nokia Technologies Oy filed Critical Nokia Technologies Oy
Publication of EP4602807A1 publication Critical patent/EP4602807A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/117Filters, e.g. for pre-processing or post-processing
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/70Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards

Definitions

  • the examples and non-limiting embodiments relate generally to multimedia transport and neural networks, and more particularly, to method, apparatus, and computer program product for integrating neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF).
  • SEI neural-network post-filter supplemental enhancement information
  • ISOBMFF ISO base media file format
  • An example apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: store coded media data in a media track of a file; store information in the media track or in a second track associated with the media track; and indicate, in the file, the presence of the information in the media track or the second track with at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information.
  • the example apparatus may further include, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information SEI.
  • NPF neural network post filter
  • the example apparatus may further include, wherein the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPF A) SEI message.
  • NNPF neural network post filter
  • SEI supplemental enhancement information
  • the example apparatus may further include, wherein, the coded media data comprise coded video data, media track comprises a video track, a restricted media sample entry type comprises a restricted video sample entry type, and the second video track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track.
  • VVC versatile video coding
  • non-VCL non-video coding layer
  • the example apparatus may further include, wherein the scheme type of a restricted media sample entry type comprises a restricted media scheme type, and wherein the apparatus is further caused to define the restricted media scheme type to: indicate that the media track requires handing of the NNPFC SEI message and/or the NNPFA SEI message for post-processing; and entail limitations on the NNPFC message and/or the NNPFA SEI message.
  • the example apparatus may further include, wherein: the restricted media scheme type comprises a first specific scheme type to indicate that the media track complies with a specific scheme and requires to process the NNPFC SEI message and the NNPFA SEI message for post-processing; the restricted media scheme type comprises a second specific scheme type to indicate that the media track complies with the specific scheme and requires the processing of the NNPFC SEI message and the NNPFA SEI messages for post-processing, wherein the NNPFC SEI message with a filter mode equal to 0 indicates base post-processing filters being used for decoding or post-processing, and wherein the NNPFC SEI message with the filter mode equal to 1 indicates filter updates being used for decoding or post-processing; the restricted media scheme type comprises a third specific scheme type to indicate that the media track complies with the specific scheme and requires the processing of the NNPFC SEI message, with the filter mode equal to 1 to indicate a specific bitstream is comprised in the SEI message, and the NNPFA SEI message for post-
  • the example apparatus may further include, wherein the apparatus is further caused to define a filter information box to carry syntax elements of the NNPFC SEI message.
  • the example apparatus may further include, wherein the apparatus is further caused to define an information prefix indication box to include an information prefix indication information message to carry one or more information prefix indications for information messages of a particular value of information payload type, and wherein each information prefix indication for the information message of a particular value of payload type indicates that one or more information messages of this value of payload type are present in a coded media sequence.
  • the example apparatus may further include, wherein the apparatus is further caused to define a first sample group, and wherein when the first sample group is an essential group, a player is required to process the first sample group, and wherein the first sample group description entry comprises the NNPFC SEI message and a first grouping type parameter, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filer, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box comprises an NNPFC identifier equal to a filter identifier.
  • the example method may further include defining a first sample group, and wherein when the first sample group is an essential group, a player is required to process the first sample group, and wherein the first sample group description entry comprises the NNPFC SEI message and a first grouping type parameter, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘ 1 ’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filer, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box comprises an NNPFC identifier equal to a filter identifier.
  • the example method may further include processing a second sample group, the second sample group description entry comprises a prefix SEI NAL unit comprising the NNPFA SEI message.
  • the example method may further include, wherein processing the second sample group comprises: ignoring and skipping the VVC non-VCL track in a bitstream reconstruction, when the second sample group is an essential sample group that is present in the VVC non-VCL NAL unit and an apparatus does not recognize the first sample group; or performing following insertion of prefix SEI NAL units when the second sample group is an essential sample group that is present in the VVC non- VCL NAL unit and the apparatus recognizes the first sample group, the following insertion is optional when the second sample group is a non-essential group; when a sample is mapped to at least one nnpfa entry, the sample implicitly comprises a prefix SEI NAL unit for each layer comprised in the media track, and the prefix SEI NAL unit comprises the NNPFA SEI message from the nnpfa entry.
  • Yet another apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: store coded media data in a media track of a file; store information in the media track or in a second track associated with the media track; and indicate, in the file, the presence of the information in the media track or the second track with a sample group for the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
  • NPF neural network post filter
  • SEI supplemental enhancement information
  • the example apparatus may further include, wherein the apparatus is further caused to define a sample group description entry to carry syntax elements of the NNPFC SEI message.
  • the example apparatus may further include, wherein the apparatus is further caused to define a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box comprises an NNPFC identifier equal to a filter identifier.
  • the apparatus is further caused to define a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides
  • the example apparatus may further include, wherein the NNPFA SEI message comprises an NNPFA identifier syntax element to identify a post-processing filter that the NNPFA SEI message is associated with.
  • the example apparatus may further include, wherein an applicable post-processing filter with an NNPFC identifier equal to a NNPFA identifier is used to filter an image/picture comprising the NNPFA SEI message.
  • the example apparatus may further include, wherein the apparatus is further caused to indicate the sample group as an essential sample group and the player is required to process the information when the sample group is indicated as the essential sample group.
  • the example apparatus may further include, wherein the apparatus is further caused to include a content of SEI manifest SEI message in a codecs multipurpose internet mail extension (MIME) parameter, wherein the codecs or a part of the codecs comprises key-value pairs, wherein a key of the key-value pair indicates SEI manifest SEI message, and a respective value comprises representation of a payload of the SEI manifest SEI message.
  • MIME multipurpose internet mail extension
  • the example apparatus may further include, wherein: the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPF A) SEI message.
  • NNPFC neural network post filter characteristics
  • NPF A neural network post filter activation
  • the example apparatus may further include, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track.
  • VVC versatile video coding
  • non-VCL non-video coding layer
  • the example apparatus may further include, wherein the apparatus is caused to process the sample group, wherein the first sample group description entry comprises the NNPFC SEI message, and the apparatus is caused to read a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘ 1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the sample group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box include an NNPFC identifier equal to filter identifier.
  • the example apparatus may further include, wherein the media track or the second track comprises at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information.
  • Yet another example method includes: storing coded media data in a media track of a file; storing information in the media track or in a second track associated with the media track; and indicating, in the file, the presence of the information in the media track or the second track with a sample group for the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
  • NPF neural network post filter
  • SEI supplemental enhancement information
  • the example method may further include, wherein the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPF A) SEI message.
  • NNPF neural network post filter
  • SEI supplemental enhancement information
  • the example method may further include, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non- video coding layer (non-VCL) track.
  • VVC versatile video coding
  • the example method may further include defining a sample group description entry to carry syntax elements of the NNPFC SEI message.
  • the example method may further include, defining a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to T’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box comprises an NNPFC identifier equal to a filter identifier.
  • the example method may further include, wherein the NNPFA SEI message comprises an NNPFA identifier syntax element to identify a post-processing filter that the NNPFA SEI message is associated with.
  • the example method may further include, wherein an applicable post-processing filter with an NNPFC identifier equal to a NNPFA identifier is used to filter an image/picture comprising the NNPFA SEI message.
  • the example method may further include, indicating the sample group as an essential sample group and the player is required to process the information when the sample group is indicated as the essential sample group.
  • the example method may further include, including a content of SEI manifest SEI message in a codecs multipurpose internet mail extension (MIME) parameter, wherein the codecs or a part of the codecs comprises key -value pairs, wherein a key of the key-value pair indicates SEI manifest SEI message, and a respective value comprises representation of a payload of the SEI manifest SEI message.
  • MIME multipurpose internet mail extension
  • Still another example method includes: reading, from a sample group of information in a file, presence of the information in a media track or a second track; processing the information in the media track or the second track, in response to supporting the information; decoding a coded media data from the media track; and filtering the decoded media data with a post-filter defined by the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
  • NPF neural network post filter
  • SEI supplemental enhancement information
  • the example method may further include, wherein: the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPFA) SEI message.
  • NNPFC neural network post filter characteristics
  • NNPFA neural network post filter activation
  • the example method may further include, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non- video coding layer (non-VCL) track.
  • VVC versatile video coding
  • the example method may further include: processing the sample group, wherein the first sample group description entry comprises the NNPFC SEI message; and read a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the sample group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box include an NNPFC identifier equal to filter identifier.
  • the example method may further include, wherein to process the first sample group, the method further comprises: ignoring and skip the VVC non-VCL track in a bitstream reconstruction, when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and an apparatus does not recognize the first sample group; or performing following insertion of prefix SEI NAL units when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus recognizes the first sample group, and wherein following insertion is optional when the first sample group is a non-essential group; when a sample is mapped to at least one sample group description entry of the sample group and the sample is either a sync sample or the first sample of a sequence of samples associated with the same sample entry, the sample implicitly comprises a prefix SEI NAL unit for each layer comprised in the media track and each filter identifier value mapped to the sample, and wherein the prefix SEI NAL unit comprises the NNPFC SEI message from the sample group description entry
  • the example method may further include, wherein the media track or the second track comprises at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information.
  • the example computer-readable medium may include, wherein the computer-readable medium comprises a non-transitory computer-readable medium.
  • FIG. 1 shows schematically an electronic device employing embodiments of the examples described herein.
  • FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.
  • FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.
  • FIG. 4 shows schematically a block diagram of an encoder on a general level.
  • FIG. 5 is a block diagram showing an interface between an encoder and a decoder in accordance with the examples described herein.
  • FIG. 6 illustrates a system configured to support streaming of media data from a source to a client device.
  • FIG. 7 is a block diagram of an apparatus that may be specifically configured in accordance with an example embodiment.
  • FIG. 8 illustrates example structure of a neural network representations (NNR) bitstream.
  • FIGs. 9a, 9b, and 9c illustrate an example of the proposed sample groups, in accordance with an embodiment.
  • FIG. 10 is an example apparatus, which may be implemented in hardware, caused to perform integration of neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF), based on the examples described herein.
  • SEI neural-network post-filter supplemental enhancement information
  • ISOBMFF ISO base media file format
  • FIG. 11 illustrates an example method for writing a file, in accordance with an embodiment.
  • FIG. 12 illustrates an example method parsing a file, in accordance with an embodiment.
  • FIG. 13 illustrates another example method for writing a file, in accordance with an embodiment.
  • FIG. 14 illustrates another example method parsing a file, in accordance with an embodiment.
  • ALF adaptive loop filtering a.k.a. also known as
  • DU distributed unit eNB or eNodeB evolved Node B (for example, an LTE base station)
  • eNB or eNodeB evolved Node B (for example, an LTE base station)
  • EN-DC E-UTRA-NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN-DC
  • E-UTRA evolved universal terrestrial radio access, for example, the LTE radio access technology
  • FDC finetuning-driving content gNB (or gNodeB) base station for 5G/NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GC
  • circuitry refers to (a) hardware-only circuit implementations (e.g., implementations in analog circuitry and/or digital circuitry); (b) combinations of circuits and computer program product(s) comprising software and/or firmware instructions stored on one or more computer readable memories that work together to cause an apparatus to perform one or more functions described herein; and (c) circuits, such as, for example, a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation even when the software or firmware is not physically present.
  • This definition of ‘circuitry’ applies to all uses of this term herein, including in any claims.
  • circuitry also includes an implementation comprising one or more processors and/or portion(s) thereof and accompanying software and/or firmware.
  • circuitry as used herein also includes, for example, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, other network device, and/or other computing device.
  • the apparatus 50 may comprise a controller 56, a processor or a processor circuitry for controlling the apparatus 50.
  • the controller 56 may be connected to a memory 58 which in embodiments of the examples described herein may store both data in the form of an image, audio data and video data, and/or may also store instructions for implementation on the controller 56.
  • the controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and/or decoding of audio, image and/or video data or assisting in coding and/or decoding carried out by the controller.
  • the apparatus 50 may further comprise a card reader 48 and a smart card 46, for example, a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.
  • a card reader 48 and a smart card 46 for example, a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.
  • the apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals, for example, for communication with a cellular communications network, a wireless communications system or a wireless local area network.
  • the apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and/or for receiving radio frequency signals from other apparatus(es).
  • the apparatus 50 may comprise a camera 42 capable of recording or detecting individual frames which are then passed to the codec circuitry 54 or the controller for processing.
  • the apparatus may receive the video image data for processing from another device prior to transmission and/or storage.
  • the apparatus 50 may also receive either wirelessly or by a wired connection the image for coding/decoding.
  • the structural elements of apparatus 50 described above represent examples of means for performing a corresponding function.
  • the system 10 comprises multiple communication devices which can communicate through one or more networks.
  • the system 10 may comprise any combination of wired or wireless networks including, but not limited to, a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network, and the like), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth® personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.
  • a wireless cellular telephone network such as a GSM, UMTS, CDMA, LTE, 4G, 5G network, and the like
  • WLAN wireless local area network
  • the example communication devices shown in the system 10 may include, but are not limited to, an electronic device or apparatus 50, a combination of a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22.
  • PDA personal digital assistant
  • IMD integrated messaging device
  • the apparatus 50 may be stationary or mobile when carried by an individual who is moving.
  • the apparatus 50 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.
  • the communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocolinternet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology.
  • CDMA code division multiple access
  • GSM global systems for mobile communications
  • UMTS universal mobile telecommunications system
  • TDMA time divisional multiple access
  • FDMA frequency division multiple access
  • TCP-IP transmission control protocolinternet protocol
  • SMS short messaging service
  • MMS multimedia messaging service
  • email instant messaging service
  • IMS instant messaging service
  • Bluetooth IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology.
  • a communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to
  • a channel may refer either to a physical channel or to a logical channel.
  • a physical channel may refer to a physical transmission medium such as a wire
  • a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels.
  • a channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.
  • the embodiments may also be implemented in internet of things (loT) devices.
  • the loT may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure.
  • the convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home/building automation, and the like, to be included in the loT.
  • the loT devices are provided with an IP address as a unique identifier.
  • the loT devices may be provided with a radio transmitter, such as WLAN or Bluetooth transmitter or a RFID tag.
  • the loT devices may have access to an IP-based network via a wired network, such as an Ethernet-based network or a powerline connection (PLC).
  • PLC powerline connection
  • Typical hybrid video encoders for example, many encoder implementations of ITU-T H.264, encode the video information in two phases. Firstly pixel values in a certain picture area (or ‘block’) are predicted, for example, by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, for example, the difference between the predicted block of pixels and the original block of pixels, is coded.
  • One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently when they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
  • Scalable video coding refers to coding structure where one bitstream can include multiple representations of the content e.g. at different bitrates, resolutions or frame rates.
  • the receiver can extract the desired representation depending on its characteristics (e.g. resolution that matches best the display device).
  • a server or a network element can extract the portions of the bitstream to be transmitted to the receiver depending on e.g. the network characteristics or processing capabilities of the receiver.
  • Scalable video coding may be realized through multi-layered coding.
  • Multi-layered coding is a concept wherein an un-encoded visual representation of a scene is, by processes such as transformation and filtering, mapped into multiple dependent or independent representations (called layers).
  • One or more encoders are used to encode a layered visual representation. When the layers include redundancies, the use of a single encoder may, by using inter-layer prediction techniques, encode with a significant gain in coding efficiency.
  • Layered video coding is typically used to provide some form of scalability in services, e.g., quality scalability, spatial scalability, temporal scalability, and view scalability.
  • a portion of a scalable video bitstream that provides a certain decoded representation, such as a base quality video or a depth map video for a bitstream that also contains texture video, and is independently decodable from other portions of the scalable video bitstream, may be referred to as an independent layer.
  • a scalable video bitstream may comprise multiple independent layers, e.g., a texture video layer, a depth video layer, and an alpha map video layer.
  • a portion of a scalable video bitstream that provides a certain decoded representation or enhancement, such as a quality enhancement to a particular fidelity or a resolution enhancement to a certain picture width and height in samples, and requires decoding of one or more other layers (a.k.a. reference layers) in the scalable video bitstream due to inter-layer prediction may be referred to as a dependent layer or a predicted layer.
  • a scalable bitstream includes a "base layer", which may provide a basic representation, such as the lowest quality video available, and one or more enhancement layers.
  • the coded representation of that layer may depend on one or more of the lower layers, e.g., inter-layer prediction may be applied.
  • the motion and mode information of the enhancement layer may be predicted from lower layers.
  • the pixel data of the lower layers may be used to create prediction for the enhancement layer.
  • enhancement layer may refer to enhancing one or more aspects of reference layer(s), such as quality or resolution.
  • a portion of the bitstream that remains after removal of all enhancement layers may be referred to as the base layer.
  • layer may be conceptual, e.g., the bitstream syntax might not include signaling of layers or the signaling of layers is not in use in a scalable bitstream that conceptually comprises several layers.
  • the term scalability layer may be used interchangeably with the term layer.
  • FIG. 4 shows a block diagram of a general structure of a video encoder.
  • FIG. 4 presents an encoder for two layers, but it would be appreciated that presented encoder could be similarly extended to encode more than two layers.
  • FIG. 4 illustrates a video encoder comprising a first encoder section 500 for a base layer and a second encoder section 502 for an enhancement layer.
  • Each of the first encoder section 500 and the second encoder section 502 may comprise similar elements for encoding incoming pictures.
  • the encoder sections 500, 502 may comprise a pixel predictor 302, 402, prediction error encoder 303, 403 and prediction error decoder 304, 404.
  • FIG. 4 shows a block diagram of a general structure of a video encoder.
  • FIG. 4 presents an encoder for two layers, but it would be appreciated that presented encoder could be similarly extended to encode more than two layers.
  • FIG. 4 illustrates a video encoder comprising a first encoder section 500 for a base layer and a second encoder section
  • the pixel predictor 302, 402 also shows an embodiment of the pixel predictor 302, 402 as comprising an inter-predictor 306, 406, an intra-predictor 308, 408, a mode selector 310, 410, a filter 316, 416, and a reference frame memory 318, 418.
  • the pixel predictor 302 of the first encoder section 500 receives base layer picture(s)/image(s) 300 of a video stream to be encoded at both the inter-predictor 306 (which determines the difference between the image and a motion compensated reference frame) and the intra-predictor 308 (which determines a prediction for an image block based only on the already processed parts of current frame or picture).
  • the output of both the inter-predictor and the intra-predictor are passed to the mode selector 310.
  • the intra-predictor 308 may have more than one intra-prediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 310.
  • the mode selector 310 also receives a copy of the base layer image(s) 300.
  • the pixel predictor 402 of the second encoder section 502 receives enhancement layer picture(s)/images(s) 400 of a video stream to be encoded at both the interpredictor 406 (which determines the difference between the image and a motion compensated reference frame) and the intra-predictor 408 (which determines a prediction for an image block based only on the already processed parts of current frame or picture).
  • the output of both the inter-predictor and the intra- predictor are passed to the mode selector 410.
  • the intra-predictor 408 may have more than one intraprediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 410.
  • the mode selector 410 also receives a copy of the enhancement layer pictures 400.
  • the pixel predictor 302, 402 further receives from a preliminary reconstructor or second summing device 339, 439 the combination of the prediction representation of the image block 312, 412 and the output 338, 438 of the prediction error decoder 304, 404.
  • the preliminary reconstructed image 314, 414 may be passed to the intra-predictor 308, 408 and to the filter 316, 416.
  • the filter 316, 416 receiving the preliminary representation may filter the preliminary representation and output a final reconstructed image 340, 440 which may be saved in the reference frame memory 318, 418.
  • the reference frame memory 318 may be connected to the inter-predictor 306 to be used as the reference image against which a future base layer image 300 is compared in inter-prediction operations.
  • the reference frame memory 318 may also be connected to the inter -predictor 406 to be used as the reference image against which a future enhancement layer image(s) 400 is compared in inter-prediction operations. Moreover, the reference frame memory 418 may be connected to the inter-predictor 406 to be used as the reference image against which the future enhancement layer image(s) 400 is compared in inter -prediction operations.
  • Filtering parameters from the filter 316 of the first encoder section 500 may be provided to the second encoder section 502 subject to the base layer being selected and indicated to be source for predicting the filtering parameters of the enhancement layer according to some embodiments.
  • the prediction error encoder 303, 403 comprises a transform unit 342, 442 and a quantizer 344, 444.
  • the transform unit 342, 442 transforms the first prediction error signal 320, 420 to a transform domain.
  • the transform is, for example, the DCT transform.
  • the quantizer 344, 444 quantizes the transform domain signal, for example, the DCT coefficients, to form quantized coefficients.
  • the prediction error decoder 304, 404 receives the output from the prediction error encoder 303, 403 and performs the opposite processes of the prediction error encoder 303, 403 to produce a decoded prediction error signal 338, 438 which, when combined with the prediction representation of the image block 312, 412 at the second summing device 339, 439, produces the preliminary reconstructed image 314, 414.
  • the prediction error decoder may be considered to comprise a dequantizer 346, 446, which dequantizes the quantized coefficient values, for example, DCT coefficients, to reconstruct the transform signal and an inverse transformation unit 348, 448, which performs the inverse transformation to the reconstructed transform signal wherein the output of the inverse transformation unit 348, 448 includes reconstructed block(s).
  • the prediction error decoder may also comprise a block filter which may filter the reconstructed block(s) according to further decoded information and filter parameters.
  • the entropy encoder 330, 430 receives the output of the prediction error encoder 303, 403 and may perform a suitable entropy encoding/variable length encoding on the signal to provide a compressed signal.
  • the outputs of the entropy encoders 330, 430 may be inserted into a bitstream, for example, by a multiplexer 508.
  • FIG. 5 is a block diagram showing the interface between an encoder 501 implementing neural network based encoding 503, and a decoder 504 implementing neural network based decoding 505 in accordance with the examples described herein.
  • the encoder 501 may embody a device, a software method or a hardware circuit.
  • the encoder 501 has the goal of compressing an input data 511 (for example, an input video) to a compressed data 512 (for example, a bitstream) such that the bitrate measuring the size of compressed data 512 is minimized, and the accuracy of an analysis or processing algorithm is maximized.
  • the encoder 501 uses an encoder or compression algorithm, for example to perform neural network based encoding 503, e.g., encoding the input data by using one or more neural networks.
  • ISO/IEC 14496-15 specifies a restricted media scheme for SEI-message-based postprocessing, labelled as 'aSEI'.
  • 'aSEI' a restricted media scheme for SEI-message-based postprocessing
  • a file author may list occurring SEI message IDs and classify them into two categories: those that are deemed required by the file author for correct playback, and others. The occurrence of either type of SEI messages may be signalled in the SEI information box.
  • Sample groupings may be represented by two linked data structures: (1) a SampleToGroupBox ('sbgp' box) represents the assignment of samples to sample groups; and (2) a SampleGroupDescriptionBox fsgpd' box) includes sample group (description) entries for describing the properties of samples mapped to this entry. There may be multiple instances of the SampleToGroupBox and SampleGroupDescriptionBox based on different grouping criteria. These may be distinguished by a type field used to indicate the type of grouping. SampleToGroupBox may additionally comprise a grouping_type_parameter field that can be used e.g., to indicate a sub-type of the grouping.
  • An essential sample group description is a sample group description for which the version field is equal to 3, and the associated sample group is also referred to as an essential sample group.
  • An essential sample group description describes essential information for the associated samples, and parsers are not allowed to attempt to process any track for which unrecognized sample group descriptions marked as essential are present.
  • the essential descriptions hierarchy sample group fesgh') indicates the processing order of the essential sample group descriptions applying to a given sample. This sample group description is an essential sample group description and uses version 3 of the SampleGroupDescriptionBox.
  • the transformation is the first sample entry transformation.
  • sample_group_description_type The transformations given in sample_group_description_type are listed in the order in which a file reader applies each transformation: any sample processing described by a sample group of type sample_group_description_type[i] is applied before any sample processing described by a sample group of type sample_group_description_type[i+l].
  • the codecs MIME parameter may include identification of essential sample groups as follows: When a restricted media track is implied by a sample entry and the transformation type indicates an essential sample group (scheme_type equal to 'essg'), the value of the codecs MIME parameter is appended by the four-character code 'essg' followed by a star ('*'), further followed by the four-character codes listed in the essential descriptions hierarchy sample description, from the first entry up to but excluding the first occurrence of 'stsd' or 'cenc'. A star ('*') is used to separate the four-character codes listed in the essential descriptions hierarchy sample description.
  • the 'essential' MIME parameter when used, is composed of one or more comma-separated essential hierarchy descriptions.
  • Each essential hierarchy description is composed of one or more four-character code of essential sample group descriptions, separated with a dot.
  • the ‘codecs’ parameter includes description of the transformation used, the listed four-character codes are the ones listed in the essential descriptions hierarchy sample group description, in the same order, from the first code following the last occurrence of 'stsd' until the last listed code. Otherwise, the listed four-character codes are the ones listed in the essential descriptions hierarchy sample group description in the same order.
  • MPEG took 3GPP AHS Release 9 as a starting point for the MPEG DASH standard (ISO/IEC 23009-1: “Dynamic adaptive streaming over HTTP (DASH)-Part 1: Media presentation description and segment formats,” International Standard, 2nd Edition, , 2014).
  • 3GPP continued to work on adaptive HTTP streaming in communication with MPEG and published 3GP- DASH (Dynamic Adaptive Streaming over HTTP; 3GPP TS 26.247: “Transparent end-to-end packet- switched streaming Service (PSS); Progressive download and dynamic adaptive Streaming over HTTP (3GP-DASH)”.
  • MPEG DASH and 3GP-DASH are technically close to each other and may therefore be collectively referred to as DASH.
  • the multimedia content may be stored on an HTTP server and may be delivered using HTTP.
  • the content may be stored on the server in two parts: Media Presentation Description (MPD), which describes a manifest of the available content, its various alternatives, their URL addresses, and other characteristics; and segments, which include the actual multimedia bitstreams in the form of chunks, in a single file or multiple files.
  • MPD Media Presentation Description
  • the MDP provides the necessary information for clients to establish a dynamic adaptive streaming over HTTP.
  • the MPD includes information describing media presentation, such as an HTTP- uniform resource locator (URL) of each Segment to make GET Segment request.
  • URL HTTP- uniform resource locator
  • the DASH client may obtain the MPD e.g. by using HTTP, email, thumb drive, broadcast, or other transport methods.
  • the DASH client may become aware of the program timing, media-content availability, media types, resolutions, minimum and maximum bandwidths, and the existence of various encoded alternatives of multimedia components, accessibility features and required digital rights management (DRM), media-component locations on the network, and other content characteristics. Using this information, the DASH client may select the appropriate encoded alternative and start streaming the content by fetching the segments using e.g. HTTP GET requests. After appropriate buffering to allow for network throughput variations, the client may continue fetching the subsequent segments and also monitor the network bandwidth fluctuations. The client may decide how to adapt to the available bandwidth by fetching segments of different alternatives (with lower or higher bitrates) to maintain an adequate buffer.
  • DRM digital rights management
  • a media presentation includes a sequence of one or more Periods, each Period includes one or more Groups, each Group includes one or more adaptation sets, each adaptation sets includes one or more representations, each representation includes one or more segments.
  • a representation is one of the alternative choices of the media content or a subset thereof typically differing by the encoding choice, e.g. by bitrate, resolution, language, codec, etc.
  • the segment includes certain duration of media data, and metadata to decode and present the included media content.
  • a segment is identified by a URI and can typically be requested by a HTTP GET request.
  • a segment may be defined as a unit of data associated with an HTTP-URL and optionally a byte range that are specified by an MPD.
  • the DASH MPD complies with Extensible Markup Language (XML) and is therefore specified through elements and attributes as defined in XML.
  • XML Extensible Markup Language
  • Preselection may be defined as set of one or more tracks representing one version of the media presentation for simultaneous decoding or presentation.
  • ISOBMFF defines a preselection group box.
  • a track grouping of type 'pres' may indicate that a track contributes to a preselection.
  • the pair of track_group_id and track_group_type identifies a track group.
  • the tracks that include a particular TrackGroupTypeBox having the same value of track_group_id and track_group_type 'pres' belong to the same preselection group.
  • TrackGroupTypeBox with track_group_type equal to 'pres' (which may also be referred to as a PreselectionGroupBox) in a track indicates that this track contributes to a preselection.
  • All the tracks that have a track group with track_group_type equal to 'pres' and a particular value of track_group_id are part of the same preselection.
  • the particular value of track_group_id may also be referred to as the ID of the preselection. This means that a preselection is uniquely identified by the track_group_id of the track group.
  • the optionally present PreselectionProcessingBox may provide information on how to process the track including this box in the context of the preselection and relative to other tracks. Consequently, the content of the PreselectionProcessingBox may differ for each track within a preselection. Preselections made from only one track may not require any track-related processing. In this case, the PreselectionProcessingBox is typically not present in the PreselectionGroupBox.
  • preselection_processing is an instance of the PreselectionProcessingBox, providing information needed for processing the including track in the context of the preselection.
  • the PreselectionProcessingBox box may include information about how the tracks contributing to the preselection may be processed. Media type specific boxes may be used to describe further processing.
  • track_order may define the order of this track relative to other tracks in the preselection, as described below.
  • sample_merge_flag 1 may indicate that this track is enabled to be merged with another track, as described below.
  • Sample entry specific specifications may require that the tracks for a preselection be provided to the respective decoder instances in a specific order. Since other means, such as the track_id, are not reliable for this purpose, the track_order may be used to order tracks in a preselection relative to each other. A lower number may indicate that at a given time, the sample of the including track is provided to the decoder before the sample with the same given of other tracks with higher number for track_order.
  • the order of these tracks may not be relevant for the preselection, and samples may be provided to the decoder in any order.
  • a merge group may be defined as a group of tracks, sorted according to track_order, where one track with the sample_merge_flag set to 0 is followed by a group of consecutive tracks with the sample_merge_flag set to 1. All tracks of a merge group may be of the same media type, and all samples may be time-aligned.
  • the sample entry type is associated with a codec-specific process to merge samples of a preselection
  • this process may be used.
  • the tracks in the merge group are all of sample entry type of “mhm2” (MPEG-H 3D Audio)
  • the merging process is defined in ISO/IEC 23008-3:2019. Tracks in a merge group may have different sample entry types.
  • the sample entry type is not associated with a codec-specific process to merge samples of a preselection, the following process may be used.
  • Merging within the merge group may comprise forming tuples of track samples with the same time stamp across contributing tracks. The ordering of samples within the tuple may be determined by track_order.
  • each merge group may result in a separate output track conformant to a media type derived from the media types of the merged tracks. For tracks not part of a merge group.
  • Preselections may be qualified, for example, by language, kind, or media specific attributes, like audio rendering indication(s), audio interactivity, and/or channel layout(s). Attributes signaled in a PreselectionTrackGroupEntryBox may take precedence over attributes signaled in contributing tracks.
  • PreselectionTrackGroupEntryBox may describe only track groups identified by track_group_type equal to 'prse'. All preselections with at least one contributing track having the track_in_movie flag set to 1 may be qualified by PreselectionTrackGroupEntryBoxes. Otherwise, the presence of the PreselectionTrackGroupEntryBoxes may be optional. All attributes uniquely qualifying a preselection may be present in the PreselectionTrackGroupEntryBox of the preselection.
  • PreselectionTrackGroupEntryBox may be as follows: aligned(8) class PreselectionTrackGroupEntryBox extends TrackGroupEntryBox('prse', version, flags)
  • This box includes information on what experience is available when this preselection is selected.
  • Boxes suitable to describe a preselection include, but are not limited to, the following list of boxes: the audio element box; the audio element selection box; the extended language tag; the user data box; the track kind; the label box; the audio rendering indication; and/or the channel layout. If a UserDataBox is included in a PreselectionTrackGroupEntryBox, then it may not carry any of the above boxes.
  • the value of numTracks may be greater than the number of tracks containing a PreselectionGroupBox with the same track_group_id in this file when the preselection is split into multiple files.
  • preselection_tag may be a codec specific value that a playback system may provide to a decoder to uniquely identify one out of several preselections in the media.
  • selection_priority may be an integer that declares the priority of the preselection in cases where no other differentiation, such as through the media language, is possible. A lower number may indicate a higher priority.
  • segment_order may specify, when present, an order rule of segments that may be suggested to be followed for ordering received segments of the Preselection.
  • the following values may be specified with semantics according to ISO/IEC 23009-1: 0: undefined; 1: time-ordered; 2: fully-ordered. Other values may be reserved.
  • segment_order When segment_order is not present, its value may be inferred to be equal to 0.
  • the kind box might utilize the Role scheme defined in ISO/IEC 23009- 1, as it provides a commonly used scheme to describe characteristics of preselections.
  • this box may carry information about the initial experience of the preselection in the referenced tracks. The preselection experience may change during the playback of these tracks (e.g. audio language may change during playback). These changes are not subject to the information presented in this box.
  • NGA Next Generation Audio
  • Preselections define user experiences that may be selected by the DASH Client. Each Preselection may be uniquely identifiable and distinguishable, (e.g. by language). A preselection may encompass a subset of media components, such that the media components may be selected and combined into a complete experience. Preselections may be used to reference a set of Representations from multiple Adaptation Sets in order to produce a complete experience.
  • Preselections may also be used to indicate a pre-defined experience at the elementary- stream level (e.g.,. the DASH Client may select a pre-defined experience and provide the selection to the media engine). Preselections may be uniquely identified by a preselection tag. Users/Codecs using this tag functionality may be encouraged to provide more information on how tags defined in the MPD map to functionality in the specific codec.
  • the media component may be referenced by the @id of the adaptation set.
  • the adaptation set includes multiple media components multiplexed on the file- container level, then each media component may be mapped to a content component.
  • a representation may include multiple tracks, and each track may be mapped to a content component. Therefore, media components may be referenced by the @id of an adaptation set, or the @ id of a content component.
  • the @id of adaptation sets and content components may be unique within the scope of a period.
  • the main adaptation set is a representation of this adaptation set that may be needed for playback of the preselection.
  • the initialization segment of such a representation is needed for playback of the preselection.
  • the main adaptation set is the adaptation set that includes the initialization segment for the complete experience.
  • Each preselection may reference a main adaptation set and may reference zero, one, or more other adaptation sets.
  • the term "main adaptation set" is not to be confused with an adaptation set that has assigned the main role.
  • a partial adaption set is a representation of this adaptation set that may only be consumable together with the main adaptation set(s) within this preselection. Again, in particular for ISO BMFF, the initialization segment of a representation of the main adaptation set is needed for playback.
  • the @ value attribute of the descriptor may provide two fields, separated by a comma: the preselection tag, and the id of the adaptation sets or content components referenced by this preselection, for example, as a white space separated list in processing order.
  • the first id may reference the main Adaptation Set.
  • the syntax for the value attribute of the preselection descriptor may follow the PRESELECTION-DESCRIPTOR-VALUE as defined in the following ABNF notation according to IETF RFC 5234:
  • ID-LIST ID-VALUE [ WHITESPACE ID- VALUE]
  • ⁇ order - Default may specify the conformance rules for representations in adaptation sets within the preselection.
  • the preselection may follow the conformance rules for multi-segment tracks.
  • the Preselection may follow the conformance rules for time-ordered segment tracks.
  • the preselection may follow the conformance rules for fully-ordered segment tracks.
  • order in the ⁇ preselectionComponents attribute may specify the component order.
  • Role may specify information on role annotation scheme.
  • Viewpoint may specify information on viewpoint annotation scheme.
  • CommonAttributesElements may specify the common attributes and elements (e.g. attributes and elements from base type Representations aseType).
  • An example of XML syntax for a Preselection element may be as follows:
  • Conformance rules for multi-segment tracks may be specified. Where multiple adaptation sets indicate this type of ordering, each adaptation set and the included representations may follow the regular conformance rules for multi- segment tracks. No additional conformance rules may be defined for the representations in different adaptation sets within preselections.
  • Conformance rules for time-ordered segment track may be specified. Where multiple adaptation sets indicate this type of ordering, each adaptation set and the included representations may follow the conformance rules for multi-segment tracks.
  • the concatenation of the following may represent a conforming segment track that also conforms to the media type as specified in the @mimeType attribute for the representation of the main adaptation set: an initialization segment of one representation of the main adaptation set (specified by the first id in the @preselectionComponents attribute or the preselection descriptor; and media segment(s)/subsegment(s) of one representation from each adaptation set referenced in the preselection ordered by non-decreasing first decode time(s). It may be noted that this may not constrain the order of segments with the same first decode time.
  • the representations of all adaptation sets referenced by the preselection may be a segment/subsegment.
  • the concatenation of the following may represent a conforming segment track, which may also conform to the media type as specified in the @mimeType attribute for the representation of the main adaptation set: an initialization segment of one representation of the main adaptation set (e.g., specified by the first id in the @preselectionComponents attribute or the preselection descriptor); and media segment(s)/subsegment(s) of one representation from each adaptation set referenced in the preselection, ordered first by non-decreasing decode times and then by position in the list given in @preselectionComponents.
  • the representations of all adaptation sets referenced by the preselection may be segment/subsegment aligned.
  • WO2022/079545 describes an example method that includes defining a metadata box for a neural network representation (NNR) item data, wherein the NNR item data comprises an NNR bitstream; and defining an association between the NNR item data and an NNR configuration by using a configuration item property, wherein the NNR configuration item property comprises information about stored NNR item data.
  • NNR neural network representation
  • NNR item may be referred to as NNR item data in some embodiments.
  • the metadata for NNR items may be included in a meta box like the metadata for any other items.
  • NNR items may be stored in an ISOBMFF file or in an external file similarly to any other items as described earlier.
  • NNR items may be stored in the ItemDataBox of the meta box of an ISOBMFF file. Such storage could be at file level, movie box level, or track level.
  • a media track may be dependent on the non-timed neural network to process its samples. The process may be associated to visual enhancement of decoded media data, decoding of the media data in the sample, or alike.
  • NNR items may be stored in one or more media data boxes (e.g. MediaDataBox) of an ISOBMFF file, while the metadata for items may be included in a meta box, which may be stored at file level, movie box level, or track level.
  • media data boxes e.g. MediaDataBox
  • NNR neural network representation
  • NNC neural network compression
  • NNR establishes a toolbox of compression methods, specifying (where applicable) the resulting elements of the compressed bitstream. All of these tools can be applied to the compression of entire neural networks, and some of them may also be applied to the compression of differential updates of neural networks with respect to a base network. Such differential updates are for example useful when models are redistributed after fine-tuning or transfer learning, or when providing versions of a neural network with different compression ratios.
  • the support for incremental compression of updates of neural networks respective to a base model will be included in the 2 nd edition of NNR, which is currently being standardized.
  • NNR comprises the syntax format, semantics, associated decoding process requirements, parameter sparsification, parameter transformation methods, parameter quantization, entropy coding method and integration/signalling within existing exchange formats.
  • FIG. 8 illustrates example structure of a neural network representations (NNR) bitstream.
  • An NNR bitstream may conform to ISO/IEC 15938-17 (Compression of Neural Networks for Multimedia Content Description and Analysis).
  • NNR specifies a high-level bitstream syntax (HLS) for signaling compressed neural network data in a channel as a sequence of NNR Units as illustrated in FIG. 8.
  • HLS high-level bitstream syntax
  • an NNR bitstream 802 includes multiple elemental units termed NNR Units (e.g.
  • NNR units 804a, 804b, 804c,... 804n NNR units 804a, 804b, 804c,... 804n).
  • An NNR Unit (e.g. 804a) represents a basic high-level syntax structure, and includes three syntax elements: NNR Unit Size 806, NNR unit header 808, NNR unit payload 810.
  • NNR bitstream may include one or more aggregate NNR units.
  • An aggregate NNR unit for example, an aggregate NNR unit 812 may include an NNR unit size 812, an NNR unit header 816, and an NNR unit pay load 818.
  • An NNR unit pay load of aggregate NNR unit may include one or more NNR units.
  • the NNR unit payload 818 includes an NNR unit 820a, an NNR unit 820b, and an NNR 820c.
  • the NNR units 820a, 820b, and 820c may include an NNR unit size, an NNR unit header, and an NNR unit payload, each of which may include zero or more syntax elements.
  • a bitstream may be formed by concatenating several NNR Units, aggregate NNR units, or combination thereof.
  • NNR units may include different types of data.
  • the type of data that is included in the payload of an NNR Unit defines the NNR Unit’s type. This type is specified in the NNR unit header.
  • the following table specifies the NNR unit header types and their identifiers.
  • NNR unit is data structure for carrying neural network data and related metadata which is compressed or represented using this specification.
  • NNR units carry compressed or uncompressed information about neural network metadata, topology information, complete or partial layer data, filters, kernels, biases, quantization weights, tensors, or the like.
  • An NNR unit may include following data elements:
  • NNR unit size This data element signals the total byte size of the NNR Unit, including the NNR unit size.
  • NNR unit header This data element includes information about the NNR unit type and related metadata.
  • NNR unit payload This data element includes compressed or uncompressed data related to the neural network.
  • NNR bitstream is composed of a sequence of NNR Units and/or aggregate NNR units.
  • the first NNR unit in an NNR bitstream shall be an NNR start unit (e.g. NNR unit of type NNR_STR).
  • Neural Network topology information can be carried as NNR units of type NNR_TPL.
  • Compressed NN information can be carried as NNR units of type NNR_NDU.
  • Parameter sets can be carried as NNR units of type NNR_MPS and NNR_LPS.
  • An NNR bitstream is formed by serializing these units.
  • a sample according to ISO/IEC 14496-15 includes one or more length-field-delimited NAL units.
  • the length field may be referred to as NALULength or NALUnitLength.
  • the NAL units in samples do not begin with start codes, but rather the length fields are used for concluding NAL unit boundaries.
  • the scheme of length-field-delimited NAL units may also be referred to as length-prefixed NAL units.
  • a VVC subpicture is a rectangular region of one or more slices within a picture.
  • An encoder may treat the subpicture boundaries like picture boundaries and may turn off loop filtering across the subpicture boundaries.
  • VVC bitstream extraction or merging operations may be performed without modifications of the video coding layer (VCL) NAL units.
  • the subpicture identifiers (IDs) for the subpictures that are present in the bitstream may be indicated in the sequence parameter set(s) or picture parameter set(s).
  • ISO/IEC 14496-15 comprises the following types of tracks for carriage of VVC video: VVC track, VVC non- VCL track, VVC subpicture track.
  • a VVC track represents a VVC elementary stream by including NAL units in its samples and/or sample entries, and possibly by associating other VVC tracks containing other layers and/or sublayers of the VVC elementary stream through 'vvcb' entity group and the 'vopi' sample group or through the 'opeg' entity group, and possibly by referencing VVC subpicture tracks.
  • VVC track references VVC subpicture tracks, it is also referred to as a VVC base track.
  • a VVC base track may not contain VCL NAL units and may not be referred to by a VVC track through a 'vvcN' track reference.
  • VVC non-VCL track is a track that includes only non-VCL NAL units and is referred to by a VVC track through a 'vvcN' track reference.
  • a VVC subpicture track includes either of the following: a sequence of one or more VVC subpictures forming a rectangular region, or a sequence of one or more complete slices forming a rectangular region.
  • a sample of a VVC subpicture track contains either of the following: one or more complete subpictures that form a rectangular region, or one or more complete slices that form a rectangular region.
  • VVC subpicture tracks enable the storage of VVC subpictures as separate tracks so that any combination of subpictures can be streamed or decoded.
  • VVC subpicture tracks enable representing rectangular regions of the same video content at different bitrates or resolutions. Consequently, bitrate or resolution emphasis on regions can be adapted dynamically by selecting the VVC subpicture tracks that are streamed or decoded.
  • WO2022/079545 also describes storage of NNR coded data as NAL units in a track that is linked to a media track. Some examples are provided below:
  • NN tracks could have time-aligned samples with the samples of associated media tracks.
  • Global NN weight updates could be stored e.g. in a new box in the track-level, a sample entry, sample group entries, or samples. Storing them in sample entry or sample group entry may be beneficial over the other alternatives e.g. for unicast streaming services.
  • - NN weight updates may be aligned with the Group of Picture structures of media data and media track.
  • the NN data could be stored in the samples of NN track.
  • NNR coded data may be stored in conjunction with media data as follows:
  • NAL units which are NNR compressed and defined in audio/video codec scope.
  • Such NAL units could be collected in NN tracks as samples of this track then linked to the video track
  • the same NN track may be linked or referenced by multiple audio/video tracks. This could be especially useful when there are multiple representations of the same media track which can utilize the same NN in its samples.
  • such media samples may include further NN updates in their bitstreams as specially marked data structures (e.g., NN specific NAL units).
  • NAL units which are NNR compressed and defined in an audio/video codec context.
  • Such NAL units may be aggregated in NNR tracks as samples of NNR track and then linked to the media tracs that they apply. Time alignment of such media tracks may be done via the sample composition timing mechanisms.
  • WO2022/079545 describes a method including defining an NNR track sample including one or more NNR units.
  • the NNR track sample is linked to an NNR unit parameter via one or more of a sample entry, a sample group, or a non-timed NNR item data.
  • the method further includes defining a network abstraction layer (NAL) unit, which is aggregated in NNR tracks as a sample of NNR track and linked to the associated media track.
  • NAL network abstraction layer
  • the NNR sample track is may be independently decodable or dependent of previous sample for decoding.
  • the NNR sample track may be used by other media tracks for decoding media samples of the other media tracks.
  • a track referencing mechanism may be used to associate the NNR sample track with the other media tracks for decoding the media samples of the other media tracks.
  • the method may further include defining a network abstraction layer (NAL) unit, wherein the NAL unit is aggregated in NNR tracks as a sample of NNR track and linked to the associated media track.
  • NAL network abstraction layer
  • An entity makes one or more SEI prefix indications available along a video bitstream, e.g. in a media description, such as SDP or DASH MPD.
  • An SEI prefix indication may comprise a SEI prefix indication SEI message or may comprise an initial part or an entire syntax structure of one or more SEI messages, such as a post-filter related SEI messages.
  • An SEI prefix indication may be, but is not limited to, one or more of the following:
  • MIME media parameter(s) may be encapsulated in an SDP parameter or in an attribute of a streaming manifest (e.g. DASH MPD) or alike. Separate or same MIME media parameter(s) may be used for declarative bitstream properties, encoding capabilities, and/or preferences or requirements for bitstream to be decoded.
  • An attribute may for example be an attribute in DASH MPD.
  • NNPFC SEI messages can be big in size and since it is optional for players and decoders to support the NNPFC SEI message, there would need to be mechanisms to indicate the presence and/or type of NNPFC SEI messages for a track (in ISOBMFF) or for a representation (in DASH). Consequently, the player or decoder may determine whether to process a track or representation including NNPFC SEI messages.
  • SEI NAL units including NNPFC and/or NNPFA SEI messages can be stored in an ISO base media file similarly to any other NAL units.
  • such a storage has at least the following drawbacks as example:
  • NNPFC SEI messages have a scope of a coded video sequence, and therefore need to be re-included in the bitstream for each coded video sequence.
  • the current mechanisms lack the capability of avoiding transmitting NNPFC SEI messages with the same content repeatedly.
  • Loading an NN model to the inference engine may take time.
  • the current mechanisms for carrying NNPFC and NNPFA SEI messages do not facilitate a lookahead mechanism to prepare for loading the NN model.
  • a file writing method includes: storing coded video data in a video track of a file; storing neural-network post-filter (NNPF) supplemental enhancement information (SEI) in the video track or in a second track associated with the video track; indicating, in the file, the presence of the NNPF SEI in the video track or the second track with at least one of a scheme type of a restricted video sample entry type and an essential sample group for the NNPF SEI.
  • NNPF neural-network post-filter
  • SEI supplemental enhancement information
  • NNPF SEI may be defined as a collective term for NNPFC SEI message(s) and/or NNPFA SEI message(s) and/or SEI NAL unit(s) including NNPFC SEI message(s) and/or NNPFA SEI message(s).
  • a file parsing method includes: reading, from a file, the presence of the NNPF SEI in a video track or a second track from at least one of a scheme type of a restricted video sample entry type and an essential sample group for the NNPF SEI; in response to supporting NNPF SEI, processing the NNPF SEI in the video track or the second track; decoding coded video data from the video track; filtering the decoded video data with a post-filter defined by the NNPF SEI.
  • NNPFC and NNPFA SEI messages are included in a track that also includes coded video data, o There might be multiple tracks with the same coded video data. One track might not include NNPFC and NNPFA SEI messages and is intended for players and decoders that do not support NNPFC and NNPFA SEI messages. Other tracks may include or reference different post-filters.
  • NNPFC and NNPFA SEI messages are included in a track that is separate from the coded video data.
  • NNPFC and NNPFA SEI messages are included in a VVC non-VCL track.
  • Players and decoders need to access the track with the coded video data and additionally players and decoders may access the track with the NNPFC and NNPFA SEI messages.
  • a restricted media scheme type is defined to indicate that the track requires handling of NNPFC and NNPFA SEI messages for post-processing.
  • a restricted media scheme type may entail limitations on NNPFC and/or NNPFA SEI messages in the track, such as limitations of the allowed syntax element values of NNPFC SEI messages. Examples of such restricted media scheme type definitions are described below.
  • a specific scheme type such as 'nnfO', indicates that the track complies with the 'aSEI' scheme and requires the processing of NNPFC SEI messages with nnpfc_mode_idc equal to 0 for indicating base post-processing filters and nnpfc_mode_idc equal to 1 for filter updates and NNPFA SEI messages for post-processing.
  • the base filter that is provided to the decoding or post-processing may be identified, for example, as described in the embodiment ‘Item property for associating an nnpfc_id value to a neural network item in ISOBMFF’ .
  • a specific scheme type such as 'nnfl', indicates that the track complies with the 'aSEI' scheme and requires the processing of NNPFC SEI messages with nnpfc_mode_idc equal to 1 (indicating that an ISO/IEC 15398-17 bitstream is included in the SEI message) and NNPFA SEI messages for post-processing.
  • a specific scheme type such as 'nnf2', indicates that the track complies with the 'aSEI' scheme and requires the processing of NNPFC SEI messages with nnpfc_mode_idc equal to 2 for indicating base post-processing filters and nnpfc_mode_idc equal to 1 for filter updates and NNPFA SEI messages for post-processing.
  • the scheme types of a track can be included in the codecs MIME parameter as a list of four-character codes separated by a plus ('+') character.
  • the codecs MIME parameter may be used in a media presentation description to indicate the properties of a track or Representation.
  • an NNPFC SEI sample group is specified as described in the following paragraphs.
  • the NNPFC SEI sample group may be indicated to be an essential sample group. Consequently, a player is required to process the NNPFC SEI sample group. For example, a player processes the NNPFC SEI sample group for bitstream reconstruction. When an NNPFC SEI sample group is not indicated to be an essential sample group, the player may ignore processing of the NNPFC SEI sample group.
  • the SampleToGroupBox(es) for this sample group type may include grouping_type_parameter equal to nnpfc_id.
  • the NNPFC SEI message has a persistence scope of a coded layer video sequence (CLVS). However, the same post-filter may be used for many CLVSs. Since NNPFC SEI messages may include NNR bitstreams, they may be of substantial size. The NNPFC sample group avoids repetitive presence of NNPFC SEI messages in tracks and thus reduces file size. File writers may make NNPFC and NNPFA sample groups essential when the postprocessing filtering is mandatory. When NNPFC and NNPFA essential sample groups are present in a VVC non-VCL track, the processing of the entire track may be omitted when the reader is not capable of post-filtering.
  • CLVS coded layer video sequence
  • FIGs. 9a, 9b, and 9c illustrate an example of the proposed sample groups, in accordance with an embodiment.
  • two base post-processing filters (IDs 0 and 1) are defined.
  • Two updates on top of each base post-processing filter are provided, labeled as updates A and B on top of base post-processing filter 0 and updates C and D on top of base post-processing filter 1.
  • Both base post-processing filters persist for the entire duration of the track (Z samples).
  • the updates A, B, C, and D have persistence of a, b, c, and d samples, respectively.
  • Two NNPFA SEI messages are defined in the sample group description of the NNPFA sample group, one activating filter ID 0 and another activating filter ID 1.
  • the first three entries in the SampleToGroupBox of the NNPFA sample group activates filter IDs 1, 0, and 1 for x, y, and z samples, respectively, as further detailed below:
  • the first entry activates the post-processing filter that is obtained by applying the update C on top of the base post-processing filter 1.
  • the second entry activates the post-processing filter that is obtained by applying the update A on top of the base post-processing filter 0 for an initial part of the period of y samples until the filter update B takes place. After that, the second entry activates the post-processing filter obtained by applying the update B on top of the base post-processing filter 0 for the latter part of the period of y samples.
  • An NNPFC SEI message includes the nnpfc_id syntax element, which is an identifying number that may be used to identify the post-processing filter that the NNPFC SEI message concerns.
  • NNPFC SEI message identifies an applicable post-processing filter associated with the nnpfc_id value.
  • the use of applicable post-processing filters with different values of nnpfc_id for specific pictures is indicated with neural-network post-filter activation (NNPFA) SEI messages.
  • NNPFA neural-network post-filter activation
  • An NNPFC SEI message either specifies a base post-processing filter or includes a neural network update.
  • a base post-processing filter is identified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within a CLVS.
  • the applicable postprocessing filter is the same as the base post-processing filter. Otherwise, the applicable post-processing filter is obtained by applying the update provided as an ISO/IEC 15938-17 bitstream in a subsequent NNPFC SEI message on top of the base post-processing filter.
  • Instances of the SampleToGroupBox for the NNPFC sample group include grouping_type_parameter.
  • the grouping_type_parameter field is specified for the NNPFC sample group as follows:
  • filter_update_flag 1 indicates that all the sample group description entries referenced by this SampleToGroupBox include an NNPFC SEI message that has nnpfc_mode_idc equal to 1 and provides an update on top of abase post-processing filter.
  • filter_update_flag 0 indicates that the sample group description entries referenced by this SampleToGroupBox includes an NNPFC SEI message that specifies a base post-processing filter.
  • filter_id indicates that the sample group description entries referenced by this SampleToGroupBox include an NNPFC SEI message that has nnpfc_id equal to filter_id.
  • the post-processing filters for different nnpfc_id values are specified in different instances of the SampleToGroupBox. Further, one SampleToGroupBox specifies the base post-processing filter(s) for a particular nnpfc_id value, while another SampleToGroupBox, when any, specifies the filter updates for the same nnpfc_id value. It is therefore possible to indicate that the base post-processing filter persists over a longer period than any of the filter updates.
  • the NNPFC sample group is an essential sample group that is present in a VVC non-VCL NAL unit and a reader does not recognize the NNPFC sample group, the reader ignores and skips the VVC non-VCL track in the bitstream reconstruction.
  • nnpfc_sei_data_byte[] is a byte array that includes exactly one complete NNPFC SEI message.
  • an NNPFC SEI sample group description entry includes a prefix SEI NAL unit that includes one or more NNPFC SEI messages with the same nnpfc_id value.
  • a prefix SEI NAL unit with NNPFC SEI message(s) from a sample group description entry is implicitly included in the first sample of each run of samples mapped to this sample group description entry and in each sync sample within each run of samples mapped to this sample group description entry.
  • grouping_type_parameter includes the information whether the SampleToGroupBox includes NNPFC SEI messages defining the base post-processing filter or NNPFC SEI message defining updates of the base post-processing filter. For example, the most significant bit of grouping_type_parameter may be used for this purpose, while the remaining bits define the nnpfc_id value.
  • Different grouping type four-character codes may be used for the base post-processing filter and updates, such as 'ncbf and 'ncfu', respectively.
  • a base post-processing filter is provided through means other than a sample group description entry, including, but not limited to, any of the following:
  • An item including a neural network
  • a box directly or indirectly included in a sample entry A box directly or indirectly included in a sample entry.
  • the NNPFA SEI sample group may be indicated to be an essential sample group. Consequently, a player is required to process the NNPFA SEI sample group.
  • grouping_type_parameter is absent for the SampleToGroupBox of a NNPFA SEI sample group.
  • grouping_type_parameter indicates the layer or layers similarly to what is done for other sample groups, such as 'sap '.
  • the syntax of the sample group description entry may be referred to as NnpfaSeiEntry and may be specified as follows: class NnpfaSeiEntryO extends VisualSampleGroupEntry fnfas') ⁇ unsigned int(8) nnpfa_sei_data_byte[];
  • the byte array nnpfa_sei_data_byte[ ] includes an NNPFA SEI message.
  • the NNPFA SEI message enables to indicate efficiently the same NNPFA SEI message being used for multiple samples, since a single NNPFA SEI message may be mapped to a run of samples in a SampleToGroupBox.
  • An NNPFA SEI message includes the nnpfa_id syntax element, which is an identifying number that may be used to identify the post-processing filter that the NNPFA SEI message concerns.
  • An NNPFA SEI message indicates that the applicable post-processing filter with nnpfc_id equal to nnpfa_id may be used to filter the picture including the NNPFA SEI message.
  • the NNPFA sample group is an essential sample group that is present in a VVC non-VCL NAL unit and a reader does not recognize the NNPFA sample group, the reader ignores and skips the VVC non-VCL track in the bitstream reconstruction.
  • the sample When a sample is mapped to at least one NnpfaSeiEntry, the sample implicitly includes a prefix SEI NAL unit for each layer included in the track, and the prefix SEI NAL unit includes the NNPFA SEI message from the NnpfaSeiEntry.
  • NNPFC sample group is an essential sample group and an NNPFA sample group is present in the same track
  • the NNPFA sample group is an essential sample group and the 'esgh' sample group lists 'ncfs' and 'ncfa' in subsequent entries of the sample_group_description_type array.
  • Syntax aligned(8) class NnpfaSeiEntryO extends VisualSampleGroupEntry('nfas')
  • Semantics nnpfa_sei_data_byte[] is a byte array that includes exactly one complete NNPFA SEI message.
  • an NNPFA SEI sample group description entry includes a prefix SEI NAL unit that includes one NNPFA SEI message.
  • the NNPFC SEI and the NNPFA SEI NAL units may be signaled/specified together in the same sample group as described in the following paragraphs.
  • the SampleToGroupBox(es) for this sample group type may include grouping_type_parameter equal to nnpfc_id.
  • grouping_type_parameter is absent for the SampleToGroupBox of a NNPF SEI sample group.
  • grouping_type_parameter indicates the layer or layers similarly to what is done for other sample groups, such as 'sap '.
  • nal_unit_type indicates the type of the NAL units in the sample group; it takes a value as defined in ISO/IEC 23090-3; it is restricted to take one of the values indicating a prefix SEI NAL unit (e.g., NNPFC SEI and NNPFA SEI).
  • num_nalus indicates the number of NAL units included in the sample group.
  • nal_unit_length indicates the length in bytes of the NAL unit.
  • nal_unit includes a declarative SEI NAL unit (e.g., NNPFC SEI and NNPFA SEI), as specified in ISO/IEC 23090-3.
  • NNPF SEI sample group where a sample group description entry may be referred to as NnpfSeiEntry and includes both NNPFC SEI message and the corresponding NNPFA SEI message.
  • the NNPFA SEI message enables to indicate efficiently the same NNPFA SEI message being used for multiple samples, since a single NNPFA SEI message may be mapped to a run of samples in a SampleToGroupBox.
  • An NNPFA SEI message includes the nnpfa_id syntax element, which is an identifying number that may be used to identify the post-processing filter that the NNPFA SEI message concerns.
  • An NNPFA SEI message indicates that the applicable post-processing filter with nnpfc_id equal to nnpfa_id may be used to filter the picture including the NNPFA SEI message.
  • Instances of the SampleToGroupBox for the NNPF sample group include grouping_type_parameter.
  • the grouping_type_parameter field is specified for the NNPF sample group as follows: ⁇ unsigned int(l) filter_update_flag; unsigned int(31) filter_id;
  • filter_update_flag 1 indicates that all the sample group description entries referenced by this SampleToGroupBox include an NNPFC SEI message that has nnpfc_mode_idc equal to 1 and provides an update on top of abase post-processing filter.
  • filter_update_flag 0 indicates that the sample group description entries referenced by this SampleToGroupBox includes an NNPFC SEI message that specifies a base post-processing filter.
  • filter_id indicates that the sample group description entries referenced by this SampleToGroupBox include an NNPFC SEI message that has nnpfc_id equal to filter_id.
  • the post-processing filters for different nnpfc_id values are specified in different instances of the SampleToGroupBox. Further, one SampleToGroupBox specifies the base post-processing filter(s) for a particular nnpfc_id value, while another SampleToGroupBox, when any, specifies the filter updates for the same nnpfc_id value. It is therefore possible to indicate that the base post-processing filter persists over a longer period than any of the filter updates.
  • the NNPF sample group is an essential sample group that is present in a VVC non-VCL NAL unit and a reader does not recognize the NNPF sample group, the reader ignores and skips the VVC non-VCL track in the bitstream reconstruction.
  • an NNPFC SEI message includes a base post-processing filter and another NNPFC SEI message includes a filter update for the same nnpfc_id value and both these SEI messages are to be included in the same access unit, they are included in the same SEI NAL unit.
  • the PreselectionGroupBox may comprise neural-network post-filter information NNPFInformationBox, as described earlier.
  • the presence of this neural-network post-filter information indicates that the track includes NNPFC and NNPFA SEI messages.
  • the information may be used by a reader to select which track out of many options with neural-network post-filter information is consumed. aligned(8) class PreselectionGroupBox extends TrackGroupTypeBox('pres')
  • nnpf_sei_presence_flag 1 indicates that the preselection includes tracks storing NNPFC and NNPFA SEI messages.
  • nnpf_sei_essential_flag 1 indicates that the processing of tracks storing NNPFC and NNPFA SEI messages is essential.
  • nnpf_sei_essential_flag 0 indicates that the processing of tracks storing NNPFC and NNPFA SEI messages is optional.
  • the presence NNPFC and NNPFA SEI messages in preselection indicated in PreselectionTrackGroupEntryBox may have the following syntax aligned(8) class PreselectionTrackGroupEntryBox extends TrackGroupEntryBoxfprse', version, flags)
  • nnpf_sei_presence_flag 1 indicates that the preselection includes tracks storing NNPFC and NNPFA SEI messages.
  • nnpf_sei_essential_flag 1 indicates that the processing of tracks storing NNPFC and NNPFA SEI messages is essential.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • General Health & Medical Sciences (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • Computational Linguistics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Molecular Biology (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Physics (AREA)
  • Software Systems (AREA)
  • Health & Medical Sciences (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

Various embodiments provide an apparatus, a method, and a computer program product. An example apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: store coded media data in a media track of a file; store information in the media track or in a second track associated with the media track; and indicate, in the file, the presence of the information in the media track or the second track with a sample group for the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).

Description

APPARATUS AND METHOD FOR INTEGRATING NEURAL-NETWORK POSTFILTER SUPPLEMENTAL ENHANCEMENT INFORMATION WITH ISO BASE MEDIA
FILE FORMAT
TECHNICAL FIELD
[0001] The examples and non-limiting embodiments relate generally to multimedia transport and neural networks, and more particularly, to method, apparatus, and computer program product for integrating neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF).
BACKGROUND
[0002] It is known to provide standardized formats for exchange of neural networks.
SUMMARY
[0003] An example apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: store coded media data in a media track of a file; store information in the media track or in a second track associated with the media track; and indicate, in the file, the presence of the information in the media track or the second track with at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information.
[0004] The example apparatus may further include, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information SEI.
[0005] The example apparatus may further include, wherein the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPF A) SEI message.
[0006] The example apparatus may further include, wherein, the coded media data comprise coded video data, media track comprises a video track, a restricted media sample entry type comprises a restricted video sample entry type, and the second video track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track. [0007] The example apparatus may further include, wherein the scheme type of a restricted media sample entry type comprises a restricted media scheme type, and wherein the apparatus is further caused to define the restricted media scheme type to: indicate that the media track requires handing of the NNPFC SEI message and/or the NNPFA SEI message for post-processing; and entail limitations on the NNPFC message and/or the NNPFA SEI message.
[0008] The example apparatus may further include, wherein: the restricted media scheme type comprises a first specific scheme type to indicate that the media track complies with a specific scheme and requires to process the NNPFC SEI message and the NNPFA SEI message for post-processing; the restricted media scheme type comprises a second specific scheme type to indicate that the media track complies with the specific scheme and requires the processing of the NNPFC SEI message and the NNPFA SEI messages for post-processing, wherein the NNPFC SEI message with a filter mode equal to 0 indicates base post-processing filters being used for decoding or post-processing, and wherein the NNPFC SEI message with the filter mode equal to 1 indicates filter updates being used for decoding or post-processing; the restricted media scheme type comprises a third specific scheme type to indicate that the media track complies with the specific scheme and requires the processing of the NNPFC SEI message, with the filter mode equal to 1 to indicate a specific bitstream is comprised in the SEI message, and the NNPFA SEI message for post-processing; and/or the restricted media scheme type comprises a specific media to indicate that the media track complies with the scheme and requires the processing of the NNPFC SEI message with filter mode equal to 2 to indicate base post-processing.
[0009] The example apparatus may further include, wherein the apparatus is further caused to define a filter information box to carry syntax elements of the NNPFC SEI message.
[0010] The example apparatus may further include, wherein the apparatus is further caused to define an information prefix indication box to include an information prefix indication information message to carry one or more information prefix indications for information messages of a particular value of information payload type, and wherein each information prefix indication for the information message of a particular value of payload type indicates that one or more information messages of this value of payload type are present in a coded media sequence.
[0011] The example apparatus may further include, wherein the apparatus is further caused to define a first sample group, and wherein when the first sample group is an essential group, a player is required to process the first sample group, and wherein the first sample group description entry comprises the NNPFC SEI message and a first grouping type parameter, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filer, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box comprises an NNPFC identifier equal to a filter identifier.
[0012] The example apparatus may further include, wherein the NNPFA SEI message comprises an NNPFA identifier syntax element to identify a post-processing filter that the NNPFA SEI message is associated with.
[0013] The example apparatus may further include, wherein an applicable post-processing filter with an NNPFC identifier equal a NNPFA identifier is used to filter an image comprising the NNPFA SEI message.
[0014] The example apparatus may further include, wherein the apparatus is further caused to define a second sample group, the second sample group description entry comprises a prefix SEI NAL unit comprising the NNPFA SEI message, and wherein when the second sample group is indicated as an essential sample group, a player is required to process the second sample group.
[0015] The example apparatus may further include, wherein the apparatus is further caused to: specify that a second grouping type parameter is absent for a second mapping box of the second sample group; or specify that the second grouping type parameter indicates a layer or layers similar to other sample groups.
[0016] The example apparatus may further include, wherein the second sample group description entry comprises a prefix SEI NAL unit comprising the NNPFA SEI message.
[0017] The example apparatus may further include, wherein the apparatus is further caused to define an information message sample group, wherein the information message sample group comprises the information message.
[0018] The example apparatus may further include, wherein the apparatus is further caused to indicate the information message sample group as an essential sample group and the player is required to process the information message sample group when the information message sample group is indicated as the essential sample group. [0019] The example apparatus may further include, wherein the information message sample group comprises a third grouping type parameter, wherein the third grouping type parameter comprises: an information payload type to indicate payload type for the information message; a payload type specific parameter, wherein when the information message comprises NNPFC SEI, the payload type specific parameter comprises a unique value of each nnpfc identifier, and wherein the payload type specific parameter is specified via a description box comprising a mapping between values of the payload type specific parameter and nnpfc identifier.
[0020] Another example apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: read, from a file, presence of information in a media track or a second track with at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information; process the information in the media track or the second track, in response to supporting the information; decode a coded media data from the media track; and filter the decoded media data with a post-filter defined by the information.
[0021] The example apparatus may further include, wherein: the information comprises neural- network prost-filter (NNPF) supplemental enhancement information (SEI), and wherein the NNPF SEI comprises at least one of a NNPFC SEI message or a NNPFA SEI message; the media track comprises a video media track; second track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track; and/or the media data comprises video data.
[0022] The example apparatus may further include, wherein the apparatus is caused to process a first sample group, wherein the first sample group description entry comprises the NNPFC SEI message and a first grouping type parameter, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to T’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filer, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box include an NNPFC identifier equal to filter identifier.
[0023] The example apparatus may further include, wherein to process the first sample group, the apparatus is caused to: ignore and skip the VVC non-VCL track in a bitstream reconstruction, when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus does not recognize the first sample group; or perform following insertion of prefix SEI NAL units when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus recognizes the first sample group, and wherein following insertion is optional when the first sample group is a non-essential group; when a sample is mapped to at least one nnpfc entry and the sample is either a sync sample or the first sample of a sequence of samples associated with the same sample entry, the sample implicitly comprises a prefix SEI NAL unit for each layer comprised in the media track and each filter identifier value mapped to the sample, and wherein the prefix SEI NAL unit comprises the NNPFC SEI message from the nnpfc entry with a filter update flag equal to 0, followed by the NNPFC SEI message from the nnpfc entry with filter update flag equal to 1 , when present.
[0024] The example apparatus may further include, wherein the apparatus is caused to process a second sample group, the second sample group description entry comprises a prefix SEI NAL unit comprising the NNPFA SEI message.
[0025] The example apparatus may further include, wherein to process the second sample group, the apparatus is caused to: ignore and skip the VVC non-VCL track in a bitstream reconstruction, when the second sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus does not recognize the first sample group; or perform following insertion of prefix SEI NAL units when the second sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus recognizes the first sample group, the following insertion is optional when the second sample group is a non-essential group; when a sample is mapped to at least one nnpfa entry, the sample implicitly comprises a prefix SEI NAL unit for each layer comprised in the media track, and the prefix SEI NAL unit comprises the NNPFA SEI message from the nnpfa entry.
[0026] An example method includes: storing coded media data in a media track of a file; storing information in the media track or in a second track associated with the media track; and indicating, in the file, the presence of the information in the media track or the second track with at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information.
[0027] The example method may further include, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information SEI.
[0028] The example method may further include, wherein the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPFA) SEI message. [0029] The example method may further include, wherein, the coded media data comprise coded video data, media track comprises a video track, a restricted media sample entry type comprises a restricted video sample entry type, and the second video track comprises a versatile video coding (VVC) non- video coding layer (non-VCL) track.
[0030] The example method may further include, wherein the scheme type of a restricted media sample entry type comprises a restricted media scheme type, and wherein the method further comprises defining the restricted media scheme type for: indicating that the media track requires handing of the NNPFC SEI message and/or the NNPFA SEI message for post-processing; and entailing limitations on the NNPFC message and/or the NNPFA SEI message.
[0031] The example method may further include, wherein: the restricted media scheme type comprises a first specific scheme type to indicate that the media track complies with a specific scheme and requires to process the NNPFC SEI message and the NNPFA SEI message for post-processing; the restricted media scheme type comprises a second specific scheme type to indicate that the media track complies with the specific scheme and requires the processing of the NNPFC SEI message and the NNPFA SEI messages for post-processing, wherein the NNPFC SEI message with a filter mode equal to 0 indicates base post-processing filters being used for decoding or post-processing, and wherein the NNPFC SEI message with the filter mode equal to 1 indicates filter updates being used for decoding or post-processing; the restricted media scheme type comprises a third specific scheme type to indicate that the media track complies with the specific scheme and requires the processing of the NNPFC SEI message, with the filter mode equal to 1 to indicate a specific bitstream is comprised in the SEI message, and the NNPFA SEI message for post-processing; and/or the restricted media scheme type comprises a specific media to indicate that the media track complies with the scheme and requires the processing of the NNPFC SEI message with filter mode equal to 2 to indicate base post-processing.
[0032] The example method may further include defining a filter information box to carry syntax elements of the NNPFC SEI message.
[0033] The example method may further include defining an information prefix indication box to include an information prefix indication information message to carry one or more information prefix indications for information messages of a particular value of information payload type, and wherein each information prefix indication for the information message of a particular value of payload type indicates that one or more information messages of this value of payload type are present in a coded media sequence. [0034] The example method may further include defining a first sample group, and wherein when the first sample group is an essential group, a player is required to process the first sample group, and wherein the first sample group description entry comprises the NNPFC SEI message and a first grouping type parameter, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘ 1 ’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filer, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box comprises an NNPFC identifier equal to a filter identifier.
[0035] The example method may further include, wherein the NNPFA SEI message comprises an NNPFA identifier syntax element to identify a post-processing filter that the NNPFA SEI message is associated with.
[0036] The example method may further include, wherein an applicable post-processing filter with an NNPFC identifier equal a NNPFA identifier is used to filter an image comprising the NNPFA SEI message.
[0037] The example method may further include the method further comprising defining a second sample group, the second sample group description entry comprises a prefix SEI NAL unit comprising the NNPFA SEI message, and wherein when the second sample group is indicated as an essential sample group, a player is required to process the second sample group.
[0038] The example method may further include: specifying that a second grouping type parameter is absent for a second mapping box of the second sample group; or specifying that the second grouping type parameter indicates a layer or layers similar to other sample groups.
[0039] The example method may further include, wherein the second sample group description entry comprises a prefix SEI NAL unit comprising the NNPFA SEI message.
[0040] The example method may further include defining an information message sample group, wherein the information message sample group comprises the information message. [0041] The example method may further include indicating the information message sample group as an essential sample group and the player is required to process the information message sample group when the information message sample group is indicated as the essential sample group
[0042] The example method may further include, wherein the information message sample group comprises a third grouping type parameter, wherein the third grouping type parameter comprises: an information payload type to indicate payload type for the information message; a payload type specific parameter, wherein when the information message comprises NNPFC SEI, the payload type specific parameter comprises a unique value of each nnpfc identifier, and wherein the payload type specific parameter is specified via a description box comprising a mapping between values of the payload type specific parameter and nnpfc identifier.
[0043] Another example method includes: reading, from a file, presence of information in a media track or a second track with at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information; processing the information in the media track or the second track, in response to supporting the information; decoding a coded media data from the media track; and filtering the decoded media data with a post-filter defined by the information.
[0044] The example method may further include, wherein: the information comprises neural- network prost-filter (NNPF) supplemental enhancement information (SEI), and wherein the NNPF SEI comprises at least one of a NNPFC SEI message or a NNPFA SEI message; the media track comprises a video media track; second track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track; and/or the media data comprises video data.
[0045] The example method may further include processing a first sample group, wherein the first sample group description entry comprises the NNPFC SEI message and a first grouping type parameter, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘ 1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filer, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box include an NNPFC identifier equal to filter identifier.
[0046] The example method may further include, wherein processing the first sample group comprises: ignoring and skipping the VVC non-VCL track in a bitstream reconstruction, when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and an apparatus does not recognize the first sample group; or performing following insertion of prefix SEI NAL units when the first sample group is an essential sample group that is present in the VVC non- VCL NAL unit and the apparatus recognizes the first sample group, and wherein following insertion is optional when the first sample group is a non-essential group; when a sample is mapped to at least one nnpfc entry and the sample is either a sync sample or the first sample of a sequence of samples associated with the same sample entry, the sample implicitly comprises a prefix SEI NAL unit for each layer comprised in the media track and each filter identifier value mapped to the sample, and wherein the prefix SEI NAL unit comprises the NNPFC SEI message from the nnpfc entry with a filter update flag equal to 0, followed by the NNPFC SEI message from the nnpfc entry with filter update flag equal to 1 , when present.
[0047] The example method may further include processing a second sample group, the second sample group description entry comprises a prefix SEI NAL unit comprising the NNPFA SEI message.
[0048] The example method may further include, wherein processing the second sample group comprises: ignoring and skipping the VVC non-VCL track in a bitstream reconstruction, when the second sample group is an essential sample group that is present in the VVC non-VCL NAL unit and an apparatus does not recognize the first sample group; or performing following insertion of prefix SEI NAL units when the second sample group is an essential sample group that is present in the VVC non- VCL NAL unit and the apparatus recognizes the first sample group, the following insertion is optional when the second sample group is a non-essential group; when a sample is mapped to at least one nnpfa entry, the sample implicitly comprises a prefix SEI NAL unit for each layer comprised in the media track, and the prefix SEI NAL unit comprises the NNPFA SEI message from the nnpfa entry.
[0049] Yet another apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: store coded media data in a media track of a file; store information in the media track or in a second track associated with the media track; and indicate, in the file, the presence of the information in the media track or the second track with a sample group for the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
[0050] The example apparatus may further include, wherein the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPFA) SEI message. [0051] The example apparatus may further include, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track.
[0052] The example apparatus may further include, wherein the apparatus is further caused to define a sample group description entry to carry syntax elements of the NNPFC SEI message.
[0053] The example apparatus may further include, wherein the apparatus is further caused to define a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box comprises an NNPFC identifier equal to a filter identifier.
[0054] The example apparatus may further include, wherein the NNPFA SEI message comprises an NNPFA identifier syntax element to identify a post-processing filter that the NNPFA SEI message is associated with.
[0055] The example apparatus may further include, wherein an applicable post-processing filter with an NNPFC identifier equal to a NNPFA identifier is used to filter an image/picture comprising the NNPFA SEI message.
[0056] The example apparatus may further include, wherein the apparatus is further caused to indicate the sample group as an essential sample group and the player is required to process the information when the sample group is indicated as the essential sample group.
[0057] The example apparatus may further include, wherein the apparatus is further caused to include a content of SEI manifest SEI message in a codecs multipurpose internet mail extension (MIME) parameter, wherein the codecs or a part of the codecs comprises key-value pairs, wherein a key of the key-value pair indicates SEI manifest SEI message, and a respective value comprises representation of a payload of the SEI manifest SEI message.
[0058] Still another apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: read, from a sample group for information in a file, presence of the information in a media track or a second track; process the information in the media track or the second track, in response to supporting the information; decode a coded media data from the media track; and filter the decoded media data with a post-filter defined by the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
[0059] The example apparatus may further include, wherein: the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPF A) SEI message.
[0060] The example apparatus may further include, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track.
[0061] The example apparatus may further include, wherein the apparatus is caused to process the sample group, wherein the first sample group description entry comprises the NNPFC SEI message, and the apparatus is caused to read a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘ 1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the sample group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box include an NNPFC identifier equal to filter identifier.
[0062] The example apparatus may further include, wherein to process the first sample group, the apparatus is caused to: ignore and skip the VVC non-VCL track in a bitstream reconstruction, when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus does not recognize the first sample group; or perform following insertion of prefix SEI NAL units when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus recognizes the first sample group, and wherein following insertion is optional when the first sample group is a non-essential group; when a sample is mapped to at least one sample group description entry of the sample group and the sample is either a sync sample or the first sample of a sequence of samples associated with the same sample entry, the sample implicitly comprises a prefix SEI NAL unit for each layer comprised in the media track and each filter identifier value mapped to the sample, and wherein the prefix SEI NAL unit comprises the NNPFC SEI message from the sample group description entry with a filter update flag equal to 0, followed by the NNPFC SEI message from the sample group description entry with filter update flag equal to 1, when present.
[0063] The example apparatus may further include, wherein the media track or the second track comprises at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information.
[0064] Yet another example method includes: storing coded media data in a media track of a file; storing information in the media track or in a second track associated with the media track; and indicating, in the file, the presence of the information in the media track or the second track with a sample group for the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
[0065] The example method may further include, wherein the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPF A) SEI message.
[0066] The example method may further include, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non- video coding layer (non-VCL) track.
[0067] The example method may further include defining a sample group description entry to carry syntax elements of the NNPFC SEI message.
[0068] The example method may further include, defining a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to T’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box comprises an NNPFC identifier equal to a filter identifier. [0069] The example method may further include, wherein the NNPFA SEI message comprises an NNPFA identifier syntax element to identify a post-processing filter that the NNPFA SEI message is associated with.
[0070] The example method may further include, wherein an applicable post-processing filter with an NNPFC identifier equal to a NNPFA identifier is used to filter an image/picture comprising the NNPFA SEI message.
[0071] The example method may further include, indicating the sample group as an essential sample group and the player is required to process the information when the sample group is indicated as the essential sample group.
[0072] The example method may further include, including a content of SEI manifest SEI message in a codecs multipurpose internet mail extension (MIME) parameter, wherein the codecs or a part of the codecs comprises key -value pairs, wherein a key of the key-value pair indicates SEI manifest SEI message, and a respective value comprises representation of a payload of the SEI manifest SEI message.
[0073] Still another example method includes: reading, from a sample group of information in a file, presence of the information in a media track or a second track; processing the information in the media track or the second track, in response to supporting the information; decoding a coded media data from the media track; and filtering the decoded media data with a post-filter defined by the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
[0074] The example method may further include, wherein: the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPFA) SEI message.
[0075] The example method may further include, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non- video coding layer (non-VCL) track.
[0076] The example method may further include: processing the sample group, wherein the first sample group description entry comprises the NNPFC SEI message; and read a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the sample group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box include an NNPFC identifier equal to filter identifier.
[0077] The example method may further include, wherein to process the first sample group, the method further comprises: ignoring and skip the VVC non-VCL track in a bitstream reconstruction, when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and an apparatus does not recognize the first sample group; or performing following insertion of prefix SEI NAL units when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus recognizes the first sample group, and wherein following insertion is optional when the first sample group is a non-essential group; when a sample is mapped to at least one sample group description entry of the sample group and the sample is either a sync sample or the first sample of a sequence of samples associated with the same sample entry, the sample implicitly comprises a prefix SEI NAL unit for each layer comprised in the media track and each filter identifier value mapped to the sample, and wherein the prefix SEI NAL unit comprises the NNPFC SEI message from the sample group description entry with a filter update flag equal to 0, followed by the NNPFC SEI message from the sample group description entry with filter update flag equal to 1, when present.
[0078] The example method may further include, wherein the media track or the second track comprises at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information.
[0079] An example computer-readable medium encoded with instructions that, when executed by a computer, causing an apparatus to perform a method as described in any of the previous paragraphs.
[0080] The example computer-readable medium may include, wherein the computer-readable medium comprises a non-transitory computer-readable medium.
[0081] An apparatus comprising means for performing the methods as described in any of the previous paragraphs. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The foregoing aspects and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:
[0083] FIG. 1 shows schematically an electronic device employing embodiments of the examples described herein.
[0084] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.
[0085] FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.
[0086] FIG. 4 shows schematically a block diagram of an encoder on a general level.
[0087] FIG. 5 is a block diagram showing an interface between an encoder and a decoder in accordance with the examples described herein.
[0088] FIG. 6 illustrates a system configured to support streaming of media data from a source to a client device.
[0089] FIG. 7 is a block diagram of an apparatus that may be specifically configured in accordance with an example embodiment.
[0090] FIG. 8 illustrates example structure of a neural network representations (NNR) bitstream.
[0091] FIGs. 9a, 9b, and 9c illustrate an example of the proposed sample groups, in accordance with an embodiment.
[0092] FIG. 10 is an example apparatus, which may be implemented in hardware, caused to perform integration of neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF), based on the examples described herein. [0093] FIG. 11 illustrates an example method for writing a file, in accordance with an embodiment.
[0094] FIG. 12 illustrates an example method parsing a file, in accordance with an embodiment.
[0095] FIG. 13 illustrates another example method for writing a file, in accordance with an embodiment.
[0096] FIG. 14 illustrates another example method parsing a file, in accordance with an embodiment.
[0097] FIG. 15 is a block diagram of one possible and non-limiting system in which the example embodiments may be practiced.
DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS
[0098] The following acronyms and abbreviations that may be found in the specification and/or the drawing figures are defined as follows:
3GP 3GPP file format
3GPP 3rd Generation Partnership Project
3GPP TS 3GPP technical specification
4CC four character code
4G fourth generation of broadband cellular network technology
5G fifth generation cellular network technology
5GC 5G core network
ACC accuracy
AGT approximated ground truth data
Al artificial intelligence
AIoT Al-enabled loT
ALF adaptive loop filtering a.k.a. also known as
AMF access and mobility management function
APS adaptation parameter set
AVC advanced video coding bpp bits-per-pixel CABAC context-adaptive binary arithmetic coding
CDMA code-division multiple access
CE core experiment ctu coding tree unit
CU central unit
DASH dynamic adaptive streaming over HTTP
DCT discrete cosine transform
DSP digital signal processor
DSNN decoder-side NN
DU distributed unit eNB (or eNodeB) evolved Node B (for example, an LTE base station)
EN-DC E-UTRA-NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN-DC
E-UTRA evolved universal terrestrial radio access, for example, the LTE radio access technology
FDMA frequency division multiple access f(n) fixed-pattern bit string using n bits written (from left to right) with the left bit first.
Fl or Fl-C interface between CU and DU control interface
FDC finetuning-driving content gNB (or gNodeB) base station for 5G/NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GC
GSM Global System for Mobile communications
H.222.0 MPEG-2 Systems is formally known as ISO/IEC 13818-1 and as ITU-T Rec. H.222.0
H.26x family of video coding standards in the domain of the ITU-T
HLS high level syntax
HQ high-quality
IBC intra block copy
ID identifier
IEC International Electrotechnical Commission
IEEE Institute of Electrical and Electronics Engineers
I/F interface IMD integrated messaging device
IMS instant messaging service loT internet of things
IP internet protocol
[RAP intra random access point
ISO International Organization for Standardization
ISOBMFF ISO base media file format
ITU International Telecommunication Union
ITU-T ITU Telecommunication Standardization Sector
JPEG joint photographic experts group
LMCS luma mapping with chroma scaling
LPNN loss proxy NN
LQ low-quality
LTE long-term evolution
LZMA Lempel-Ziv-Markov chain compression
LZMA2 simple container format that can include both uncompressed data and LZMA data
LZO Lempel-Ziv-Oberhumer compression
LZW Lempel-Ziv-Welch compression
MAC medium access control mdat MediaDataBox
MME mobility management entity
MMS multimedia messaging service moov MovieBox
MP4 file format for MPEG-4 Part 14 files
MPEG moving picture experts group
MPEG-2 H.222/H.262 as defined by the ITU
MPEG-4 audio and video coding standard for ISO/IEC 14496
MSB most significant bit
MSE mean square error
NAL network abstraction layer
NDU NN compressed data unit ng or NG new generation ng-eNB or NG-eNB new generation eNB
NN neural network
NNEF neural network exchange format NNR neural network representation
NR new radio (5G radio)
N/W or NW network
ONNX Open Neural Network eXchange
PB protocol buffers
PC personal computer
PDA personal digital assistant
PDCP packet data convergence protocol
PHY physical layer
PID packet identifier
PLC power line communication
PNG portable network graphics
PSNR peak signal-to-noise ratio
QP quantization power
RAM random access memory
RAN radio access network
RBSP raw byte sequence payload
RD loss rate distortion loss
RFC request for comments
RFID radio frequency identification
RLC radio link control
RRC radio resource control
RRH remote radio head
RU radio unit
Rx receiver
SDAP service data adaptation protocol
SGD Stochastic Gradient Descent
SGW serving gateway
SMF session management function
SMS short messaging service
SPS sequence parameter set st(v) null-terminated string encoded as UTF-8 characters as specified in ISO/IEC 10646
SVC scalable video coding
SI interface between eNodeBs and the EPC
TCP-IP transmission control protocol-internet protocol TDMA time divisional multiple access trak TrackBox
TS transport stream
TUC technology under consideration
TV television
Tx transmitter
UE user equipment ue(v) unsigned integer Exp-Golomb-coded syntax element with the left bit first
UICC Universal Integrated Circuit Card
UMTS Universal Mobile Telecommunications System u(n) unsigned integer using n bits
UPF user plane function
URI uniform resource identifier
URL uniform resource locator
UTF-8 8-bit Unicode Transformation Format
VPS video parameter set
VVC versatile video coding
WLAN wireless local area network
X2 interconnecting interface between two eNodeBs in LTE network
Xn interface between two NG-RAN nodes
[0099] Some embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments of the invention are shown. Indeed, various embodiments of the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals refer to like elements throughout. As used herein, the terms ‘data,’ ‘content,’ ‘information,’ and similar terms may be used interchangeably to refer to data capable of being transmitted, received and/or stored in accordance with embodiments of the present invention. Thus, use of any such terms should not be taken to limit the spirit and scope of embodiments of the present invention.
[0100] Additionally, as used herein, the term ‘circuitry’ refers to (a) hardware-only circuit implementations (e.g., implementations in analog circuitry and/or digital circuitry); (b) combinations of circuits and computer program product(s) comprising software and/or firmware instructions stored on one or more computer readable memories that work together to cause an apparatus to perform one or more functions described herein; and (c) circuits, such as, for example, a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation even when the software or firmware is not physically present. This definition of ‘circuitry’ applies to all uses of this term herein, including in any claims. As a further example, as used herein, the term ‘circuitry’ also includes an implementation comprising one or more processors and/or portion(s) thereof and accompanying software and/or firmware. As another example, the term ‘circuitry’ as used herein also includes, for example, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, other network device, and/or other computing device.
[0101] As defined herein, a ‘computer-readable storage medium,’ which refers to a non-transitory physical storage medium (e.g., volatile or non-volatile memory device), can be differentiated from a ‘computer-readable transmission medium,’ which refers to an electromagnetic signal.
[0102] A method, apparatus and computer program product are provided in accordance with example embodiments for integrating neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF).
[0103] In an example, the following describes in detail suitable apparatus and possible mechanisms for integrating neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF). In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an apparatus 50. The apparatus may be an Internet of Things (loT) apparatus configured to perform various functions, for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 will be explained next.
[0104] The apparatus 50 may for example be a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or a lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.
[0105] The apparatus 50 may comprise a housing 30 for incorporating and protecting the device. The apparatus 50 may further comprise a display 32, for example, in the form of a liquid crystal display, light emitting diode display, organic light emitting diode display, and the like. In other embodiments of the examples described herein the display may be any suitable display technology suitable to display media or multimedia content, for example, an image or a video. The apparatus 50 may further comprise a keypad 34. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0106] The apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input. The apparatus 50 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 38, speaker, or an analogue audio or digital audio output connection. The apparatus 50 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera 42 capable of recording or capturing images and/or video. The apparatus 50 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 50 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB/firewire wired connection.
[0107] The apparatus 50 may comprise a controller 56, a processor or a processor circuitry for controlling the apparatus 50. The controller 56 may be connected to a memory 58 which in embodiments of the examples described herein may store both data in the form of an image, audio data and video data, and/or may also store instructions for implementation on the controller 56. The controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and/or decoding of audio, image and/or video data or assisting in coding and/or decoding carried out by the controller.
[0108] The apparatus 50 may further comprise a card reader 48 and a smart card 46, for example, a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.
[0109] The apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals, for example, for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and/or for receiving radio frequency signals from other apparatus(es). [0110] The apparatus 50 may comprise a camera 42 capable of recording or detecting individual frames which are then passed to the codec circuitry 54 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and/or storage. The apparatus 50 may also receive either wirelessly or by a wired connection the image for coding/decoding. The structural elements of apparatus 50 described above represent examples of means for performing a corresponding function.
[0111] With respect to FIG. 3, an example of a system within which embodiments of the examples described herein can be utilized is shown. The system 10 comprises multiple communication devices which can communicate through one or more networks. The system 10 may comprise any combination of wired or wireless networks including, but not limited to, a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network, and the like), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth® personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.
[0112] The system 10 may include both wired and wireless communication devices and/or apparatus 50 suitable for implementing embodiments of the examples described herein.
[0113] For example, the system shown in FIG. 3 shows a mobile telephone network 11 and a representation of the Internet 28. Connectivity to the Internet 28 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.
[0114] The example communication devices shown in the system 10 may include, but are not limited to, an electronic device or apparatus 50, a combination of a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22. The apparatus 50 may be stationary or mobile when carried by an individual who is moving. The apparatus 50 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.
[0115] The embodiments may also be implemented in a set-top box; for example, a digital TV receiver, which may/may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and/or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and/or embedded systems offering hardware/software based coding. [0116] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that allows communication between the mobile telephone network 11 and the Internet 28. The system may include additional communication devices and communication devices of various types.
[0117] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocolinternet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0118] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.
[0119] The embodiments may also be implemented in internet of things (loT) devices. The loT may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home/building automation, and the like, to be included in the loT. In order to utilize the loT devices are provided with an IP address as a unique identifier. The loT devices may be provided with a radio transmitter, such as WLAN or Bluetooth transmitter or a RFID tag. Alternatively, the loT devices may have access to an IP-based network via a wired network, such as an Ethernet-based network or a powerline connection (PLC).
[0120] The devices/systems described in FIGs. 1 to 3 enable encoding, decoding, and/or transportation of, for example, a neural network representation and/or a media bitstream. [0121] An MPEG-2 transport stream (TS), specified in ISO/IEC 13818-1 or equivalently in ITU- T Recommendation H.222.0, is a format for carrying audio, video, and other media as well as program metadata or other metadata, in a multiplexed stream. A packet identifier (PID) is used to identify an elementary stream (a.k.a. packetized elementary stream) within the TS. Hence, a logical channel within an MPEG-2 TS may be considered to correspond to a specific PID value.
[0122] Available media file format standards include ISO base media file format (ISO/IEC 14496- 12, which may be abbreviated ISOBMFF) and the file format for NAL unit structured video (ISO/IEC 14496-15), which derives from the ISOBMFF.
[0123] Video codec includes an encoder that transforms the input video into a compressed representation suited for storage/transmission and a decoder that can decompress the compressed video representation back into a viewable form, or into a form that is suitable as an input to one or more algorithms for analysis or processing. A video encoder and/or a video decoder may also be separate from each other, for example, need not form a codec. Typically, encoder discards some information in the original video sequence in order to represent the video in a more compact form (e.g., at lower bitrate).
[0124] Typical hybrid video encoders, for example, many encoder implementations of ITU-T H.264, encode the video information in two phases. Firstly pixel values in a certain picture area (or ‘block’) are predicted, for example, by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, for example, the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (for example, Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).
[0125] In temporal prediction, the sources of prediction are previously decoded pictures (a.k.a. reference pictures). In intra block copy (IBC; a.k.a. intra-block-copy prediction and current picture referencing), prediction is applied similarly to temporal prediction, but the reference picture is the current picture and only previously decoded samples can be referred in the prediction process. Interlayer or inter-view prediction may be applied similarly to temporal prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal prediction only, while in other cases inter prediction may refer collectively to temporal prediction and any of intra block copy, inter-layer prediction, and inter-view prediction provided that they are performed with the same or similar process than temporal prediction. Inter prediction or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0126] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, reduces temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures. Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, for example, either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra-coding, where no inter prediction is applied.
[0127] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently when they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0128] Scalable video coding refers to coding structure where one bitstream can include multiple representations of the content e.g. at different bitrates, resolutions or frame rates. In these cases, the receiver can extract the desired representation depending on its characteristics (e.g. resolution that matches best the display device). Alternatively, a server or a network element can extract the portions of the bitstream to be transmitted to the receiver depending on e.g. the network characteristics or processing capabilities of the receiver.
[0129] Scalable video coding may be realized through multi-layered coding. Multi-layered coding is a concept wherein an un-encoded visual representation of a scene is, by processes such as transformation and filtering, mapped into multiple dependent or independent representations (called layers). One or more encoders are used to encode a layered visual representation. When the layers include redundancies, the use of a single encoder may, by using inter-layer prediction techniques, encode with a significant gain in coding efficiency. Layered video coding is typically used to provide some form of scalability in services, e.g., quality scalability, spatial scalability, temporal scalability, and view scalability. [0130] A portion of a scalable video bitstream that provides a certain decoded representation, such as a base quality video or a depth map video for a bitstream that also contains texture video, and is independently decodable from other portions of the scalable video bitstream, may be referred to as an independent layer. A scalable video bitstream may comprise multiple independent layers, e.g., a texture video layer, a depth video layer, and an alpha map video layer. A portion of a scalable video bitstream that provides a certain decoded representation or enhancement, such as a quality enhancement to a particular fidelity or a resolution enhancement to a certain picture width and height in samples, and requires decoding of one or more other layers (a.k.a. reference layers) in the scalable video bitstream due to inter-layer prediction may be referred to as a dependent layer or a predicted layer.
[0131] In some scenarios, a scalable bitstream includes a "base layer", which may provide a basic representation, such as the lowest quality video available, and one or more enhancement layers. In order to improve coding efficiency for an enhancement layer, the coded representation of that layer may depend on one or more of the lower layers, e.g., inter-layer prediction may be applied. For example, the motion and mode information of the enhancement layer may be predicted from lower layers. Similarly, the pixel data of the lower layers may be used to create prediction for the enhancement layer. The term enhancement layer may refer to enhancing one or more aspects of reference layer(s), such as quality or resolution. A portion of the bitstream that remains after removal of all enhancement layers may be referred to as the base layer.
[0132] It needs to be understood that the term layer may be conceptual, e.g., the bitstream syntax might not include signaling of layers or the signaling of layers is not in use in a scalable bitstream that conceptually comprises several layers. The term scalability layer may be used interchangeably with the term layer.
[0133] FIG. 4 shows a block diagram of a general structure of a video encoder. FIG. 4 presents an encoder for two layers, but it would be appreciated that presented encoder could be similarly extended to encode more than two layers. FIG. 4 illustrates a video encoder comprising a first encoder section 500 for a base layer and a second encoder section 502 for an enhancement layer. Each of the first encoder section 500 and the second encoder section 502 may comprise similar elements for encoding incoming pictures. The encoder sections 500, 502 may comprise a pixel predictor 302, 402, prediction error encoder 303, 403 and prediction error decoder 304, 404. FIG. 4 also shows an embodiment of the pixel predictor 302, 402 as comprising an inter-predictor 306, 406, an intra-predictor 308, 408, a mode selector 310, 410, a filter 316, 416, and a reference frame memory 318, 418. The pixel predictor 302 of the first encoder section 500 receives base layer picture(s)/image(s) 300 of a video stream to be encoded at both the inter-predictor 306 (which determines the difference between the image and a motion compensated reference frame) and the intra-predictor 308 (which determines a prediction for an image block based only on the already processed parts of current frame or picture). The output of both the inter-predictor and the intra-predictor are passed to the mode selector 310. The intra-predictor 308 may have more than one intra-prediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 310. The mode selector 310 also receives a copy of the base layer image(s) 300. Correspondingly, the pixel predictor 402 of the second encoder section 502 receives enhancement layer picture(s)/images(s) 400 of a video stream to be encoded at both the interpredictor 406 (which determines the difference between the image and a motion compensated reference frame) and the intra-predictor 408 (which determines a prediction for an image block based only on the already processed parts of current frame or picture). The output of both the inter-predictor and the intra- predictor are passed to the mode selector 410. The intra-predictor 408 may have more than one intraprediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 410. The mode selector 410 also receives a copy of the enhancement layer pictures 400.
[0134] Depending on which encoding mode is selected to encode the current block, the output of the inter-predictor 306, 406 or the output of one of the optional intra-predictor modes or the output of a surface encoder within the mode selector is passed to the output of the mode selector 310, 410. The output of the mode selector 310, 410 is passed to a first summing device 321, 421. The first summing device may subtract the output of the pixel predictor 302, 402 from the base layer image(s) 300/enhancement layer image(s) 400 to produce a first prediction error signal 320, 420 which is input to the prediction error encoder 303, 403.
[0135] The pixel predictor 302, 402 further receives from a preliminary reconstructor or second summing device 339, 439 the combination of the prediction representation of the image block 312, 412 and the output 338, 438 of the prediction error decoder 304, 404. The preliminary reconstructed image 314, 414 may be passed to the intra-predictor 308, 408 and to the filter 316, 416. The filter 316, 416 receiving the preliminary representation may filter the preliminary representation and output a final reconstructed image 340, 440 which may be saved in the reference frame memory 318, 418. The reference frame memory 318 may be connected to the inter-predictor 306 to be used as the reference image against which a future base layer image 300 is compared in inter-prediction operations. Subject to the base layer being selected and indicated to be source for inter-layer sample prediction and/or interlayer motion information prediction of the enhancement layer according to some embodiments, the reference frame memory 318 may also be connected to the inter -predictor 406 to be used as the reference image against which a future enhancement layer image(s) 400 is compared in inter-prediction operations. Moreover, the reference frame memory 418 may be connected to the inter-predictor 406 to be used as the reference image against which the future enhancement layer image(s) 400 is compared in inter -prediction operations.
[0136] Filtering parameters from the filter 316 of the first encoder section 500 may be provided to the second encoder section 502 subject to the base layer being selected and indicated to be source for predicting the filtering parameters of the enhancement layer according to some embodiments.
[0137] The prediction error encoder 303, 403 comprises a transform unit 342, 442 and a quantizer 344, 444. The transform unit 342, 442 transforms the first prediction error signal 320, 420 to a transform domain. The transform is, for example, the DCT transform. The quantizer 344, 444 quantizes the transform domain signal, for example, the DCT coefficients, to form quantized coefficients.
[0138] The prediction error decoder 304, 404 receives the output from the prediction error encoder 303, 403 and performs the opposite processes of the prediction error encoder 303, 403 to produce a decoded prediction error signal 338, 438 which, when combined with the prediction representation of the image block 312, 412 at the second summing device 339, 439, produces the preliminary reconstructed image 314, 414. The prediction error decoder may be considered to comprise a dequantizer 346, 446, which dequantizes the quantized coefficient values, for example, DCT coefficients, to reconstruct the transform signal and an inverse transformation unit 348, 448, which performs the inverse transformation to the reconstructed transform signal wherein the output of the inverse transformation unit 348, 448 includes reconstructed block(s). The prediction error decoder may also comprise a block filter which may filter the reconstructed block(s) according to further decoded information and filter parameters.
[0139] The entropy encoder 330, 430 receives the output of the prediction error encoder 303, 403 and may perform a suitable entropy encoding/variable length encoding on the signal to provide a compressed signal. The outputs of the entropy encoders 330, 430 may be inserted into a bitstream, for example, by a multiplexer 508.
[0140] FIG. 5 is a block diagram showing the interface between an encoder 501 implementing neural network based encoding 503, and a decoder 504 implementing neural network based decoding 505 in accordance with the examples described herein. The encoder 501 may embody a device, a software method or a hardware circuit. The encoder 501 has the goal of compressing an input data 511 (for example, an input video) to a compressed data 512 (for example, a bitstream) such that the bitrate measuring the size of compressed data 512 is minimized, and the accuracy of an analysis or processing algorithm is maximized. To this end, the encoder 501 uses an encoder or compression algorithm, for example to perform neural network based encoding 503, e.g., encoding the input data by using one or more neural networks.
[0141] The general analysis or processing algorithm may be part of the decoder 504. The decoder 504 uses a decoder or decompression algorithm, for example, to perform the neural network based decoding 505 (e.g., decoding by using one or more neural networks) to decode the compressed data 512 (for example, compressed video) which was encoded by the encoder 501. The decoder 504 produces decompressed data 513 (for example, reconstructed data).
[0142] The encoder 501 and decoder 504 may be entities implementing an abstraction, may be separate entities or the same entities, or may be part of the same physical device.
[0143] The analysis/processing algorithm may be any algorithm, traditional or learned from data. In the case of an algorithm which is learned from data, in some embodiments it is assumed that this algorithm can be modified or updated, for example, by using optimization via gradient descent. An example of the learned algorithm is a neural network.
[0144] An out-of-band transmission, signaling, or storage may refer to the capability of transmitting, signaling, or storing information in a manner that associates the information with a video bitstream. The out-of-band transmission may use a more reliable transmission mechanism compared to the protocols used for carrying coded video data, such as slices. The out-of-band transmission, signaling or storage can additionally or alternatively be used e.g. for ease of access or session negotiation. For example, a sample entry of a track in a file conforming to the ISO Base Media File Format may comprise parameter sets, while the coded data in the bitstream is stored elsewhere in the file or in another file. Another example of out-of-band transmission, signaling, or storage comprises including information, such as NN and/or NN updates in a file format track that is separate from track(s) including coded video data.
[0145] The phrase along the bitstream (e.g. indicating along the bitstream) or along a coded unit of a bitstream (e.g. indicating along a coded tile) may be used in claims and described embodiments to refer to transmission, signaling, or storage in a manner that the ‘out-of-band’ data is associated with, but not included within, the bitstream or the coded unit, respectively. The phrase decoding along the bitstream or along a coded unit of a bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream or the coded unit, respectively. For example, the phrase along the bitstream may be used when the bitstream is included in a container file, such as a file conforming to the ISO Base Media File Format, and certain file metadata is stored in the file in a manner that associates the metadata to the bitstream, such as boxes in the sample entry for a track including the bitstream, a sample group for the track including the bitstream, or a timed metadata track associated with the track including the bitstream. In another example, the phrase along the bitstream may be used when the bitstream is made available as a stream over a communication protocol and a media description, such as a streaming manifest, is provided to describe the stream.
[0146] A bitstream may be defined as a sequence of bits or a sequence of syntax structures. A bitstream format may constrain the order of syntax structures in the bitstream.
[0147] A syntax element may be defined as an element of data represented in a bitstream. A syntax structure may be defined as zero or more syntax elements present together in a bitstream in a specified order.
[0148] An elementary unit for the output of a video encoder and the input of a video decoder, respectively, may be a network abstraction layer (NAL) unit. For transport over packet-oriented networks or storage into structured files, NAL units may be encapsulated into packets or similar structures. A bytestream format encapsulating NAL units may be used for transmission or storage environments that do not provide framing structures. The bytestream format may separate NAL units from each other by attaching a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, encoders may run a byte-oriented start code emulation prevention algorithm, which may add an emulation prevention byte to the NAL unit payload when a start code would have occurred otherwise. In order to enable straightforward gateway operation between packet and stream-oriented systems, start code emulation prevention may be performed regardless of whether the bytestream format is in use or not. A NAL unit may be defined as a syntax structure including an indication of the type of data to follow and bytes including that data in the form of a raw byte sequence payload interspersed as necessary with emulation prevention bytes. A raw byte sequence payload (RBSP) may be defined as a syntax structure including an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits including syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0149] A bitstream may be defined to logically include a syntax structure, such as a NAL unit, when the syntax structure is transmitted along the bitstream but may be included in the bitstream according to the bitstream format. A bitstream may be defined to natively comprise a syntax structure, when the bitstream includes the syntax structure. [0150] In some coding formats or standards, a bitstream may be in the form of a network abstraction layer (NAL) unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.
[0151] In some coding formats or standards, the end of a bitstream may be indicated by a specific NAL unit, which may be referred to as the end of bitstream (EOB) NAL unit and which is the last NAL unit of the bitstream.
[0152] In some formats or standards, a first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams.
[0153] In some coding formats, such as AVI, a bitstream may comprise a sequence of open bitstream units (OBUs). An OBU comprises a header and a payload, wherein the header identifies a type of the OBU. Furthermore, the header may comprise a size of the payload in bytes.
[0154] In some coding standards, NAL units include of a header and pay load. The NAL unit header indicates the type of the NAL unit. In some coding standards, the NAL unit header indicates a scalability layer identifier (e.g. called nuh_layer_id in H.265/HEVC and H.266/VVC), which could be used e.g. for indicating spatial or quality layers, views of a multiview video, or auxiliary layers (such as depth maps or alpha planes). In some coding standards, the NAL unit header includes a temporal sublayer identifier, which may be used for indicating temporal subsets of the bitstream, such as a 30- frames-per-second subset of a 60-frames-per-second bitstream.
[0155] NAL units may be categorized into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are typically coded slice NAL units.
[0156] A non-VCL NAL unit may be, for example, one of the following types: a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end of sequence NAL unit, an end of bitstream NAL unit, or a filler data NAL unit. Parameter sets may be needed for the reconstruction of decoded pictures, whereas many of the other non-VCL NAL units are not necessary for the reconstruction of decoded sample values.
[0157] Some coding formats specify parameter sets that may carry parameter values needed for the decoding or reconstruction of decoded pictures. A parameter may be defined as a syntax element of a parameter set. A parameter set may be defined as a syntax structure that includes parameters and that can be referred to from or activated by another syntax structure, for example, using an identifier.
[0158] Some types of parameter sets are briefly described in the following, but it needs to be understood, that other types of parameter sets may exist and that embodiments may be applied, but are not limited to, the described types of parameter sets.
[0159] Parameters that remain unchanged through a coded video sequence may be included in a sequence parameter set. Alternatively, an SPS may be limited to apply to a layer that references the SPS, e.g. an SPS may remain valid for a coded layer video sequence. In addition to the parameters that may be needed by the decoding process, the sequence parameter set may optionally include video usability information (VUI), which includes parameters that may be important for buffering, picture output timing, rendering, and resource reservation.
[0160] A picture parameter set includes such parameters that are likely to be unchanged in several coded pictures. A picture parameter set may include parameters that can be referred to by the VCL NAL units of one or more coded pictures.
[0161] A video parameter set (VPS) may be defined as a syntax structure including syntax elements that apply to zero or more entire coded video sequences and may include parameters applying to multiple layers. The VPS may provide information about the dependency relationships of the layers in a bitstream, as well as many other information that are applicable to all slices across all layers in the entire coded video sequence.
[0162] A video parameter set RBSP may include parameters that can be referred to by one or more sequence parameter set RBSPs.
[0163] The relationship and hierarchy between a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS) may be described as follows. A VPS resides one level above an SPS in the parameter set hierarchy and in the context of scalability. The VPS may include parameters that are common for all slices across all layers in the entire coded video sequence. The SPS includes the parameters that are common for all slices in a particular layer in the entire coded video sequence, and may be shared by multiple layers. The PPS includes the parameters that are common for all slices in a particular picture and are likely to be shared by all slices in multiple pictures. [0164] An adaptation parameter set (APS) may be specified in some coding formats, such as H.266/VVC. An APS may be applied to one or more image segments, such as slices. In H.266/V VC, an APS may be defined as a syntax structure including syntax elements that apply to zero or more slices as determined by zero or more syntax elements found in slice headers or in a picture header. An APS may comprise a type (aps_params_type in H.266/VVC) and an identifier (aps_adaptation_parameter_set_id in H.266/VVC). The combination of an APS type and an APS identifier may be used to identify a particular APS. H.266/VVC comprises three APS types: an adaptive loop filtering (ALF), a luma mapping with chroma scaling (LMCS), and a scaling list APS types. The ALF APS(s) are referenced from a slice header (thus, the referenced ALF APSs can change slice by slice), and the LMCS and scaling list APS(s) are referenced from a picture header (thus, the referenced LMCS and scaling list APSs can change picture by picture). In H.266/VVC, the APS RBSP has the following syntax:
[0165] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI NAL units, and some video coding specifications include both prefix SEI NAL units and suffix SEI NAL units. A prefix SEI NAL unit can start a picture unit or alike; and a suffix SEI NAL unit can end a picture unit or alike. Hereafter, an SEI NAL unit may equivalently refer to a prefix SEI NAL unit or a suffix SEI NAL unit. An SEI NAL unit includes one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation.
[0166] Several SEI messages are specified in H.264/AVC, H.265/HEVC, H.266/VVC, and H.274/VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for specific use. The standards may include the syntax and semantics for the specified SEI messages but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient can be specified.
[0167] Some video coding specifications enable metadata OBUs. A metadata OBU comprises a type field, which specifies the type of metadata.
[0168] A coded video sequence (CVS) may be defined as a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream.
[0169] A coded layer video sequence (CLVS) may be defined as a sequence of pictures and associated other data within the same scalable layer (e.g., with the same value of nuh_layer_id in VVC) that is decodable independently of other pictures in the same layer.
[0170] An identifier may be defined as a syntax element that identifies a syntax structure. A value of the identifier may for example differ in different instances of the same syntax structure, such as a parameter set. A particular instance of the syntax structure may be referenced through its identifier value. For example, a parameter set that is referenced by the (de)coding of a coded video slice may be identified by providing the identifier value of the parameter set in a header of the coded video slice.
[0171] An indicator (ide) may be defined as a syntax element whose value indicates a selection among more than two values (for which semantics have been specified). An indicator syntax element may have _idc postfix in its name. [0172] A uniform resource identifier (URI) may be defined as a string of characters used to identify a name of a resource. Such identification enables interaction with representations of the resource over a network, using specific protocols. A URI is defined through a scheme specifying a concrete syntax and associated protocol for the URI. The uniform resource locator (URL) and the uniform resource name (URN) are forms of URI. A URL may be defined as a URI that identifies a web resource and specifies the means of acting upon or obtaining the representation of the resource, specifying both its primary access mechanism and network location. A URN may be defined as a URI that identifies a resource by name in a particular namespace. A URN may be used for identifying a resource without implying its location or how to access it.
[0173] Internet media types, also known as multipurpose internet mail extension (MIME) types, are used by various applications to identify the type of a resource or a file. MIME types include a media type (e.g., ‘image’ in the case of still images), a subtype, and zero or more optional parameters.
[0174] The MIME is an extension to an email protocol which makes it possible to transmit and receive different kinds of data files on the Internet, for example video, audio, images, and software. An internet media type is an identifier used on the Internet to indicate the type of data that a file includes. Such internet media types may also be called as content types. Several MIME type/subtype combinations exist that may include different media formats. Content type information may be included by a transmitting entity in a MIME header at the beginning of a media transmission. A receiving entity thus may need to examine the details of such media content to determine when the specific elements may be rendered given an available set of codecs. Especially, when the end system has limited resources, or the connection to the end system has limited bandwidth, it may be helpful to know from the content type alone if the content can be rendered.
[0175] One of the original motivations for MIME is the ability to identify the specific media type of a message part. However, due to various factors, it is not always possible from looking at the MIME type and subtype to know which specific media formats are included in the body part or which codecs are indicated in order to render the content. Optional media parameters may be provided in addition to the MIME type and subtype to provide further details of the media content.
[0176] An optional ‘codecs’ MIME parameter is specified to be used with various MIME types or type/subtype combinations to allow for unambiguous specification of the codecs employed by the media formats included within the overall container format.
[0177] By labelling content with the specific codecs indicated to render the included media, receiving systems may determine when the codecs are supported by the end system, and when not, may take appropriate action (such as rejecting the content, sending notification of the situation, transcoding the content to a supported type, fetching and installing the required codecs, further inspection to determine when it may be sufficient to support a subset of the indicated codecs, and the like).
[0178] For file formats derived from the ISOBMFF, the codecs parameter may be considered to comprise a comma-separated list of one or more list items.
[0179] When a list item of the codecs parameter represents a track of an ISOBMFF compliant file, the list item may comprise a four-character code of the sample entry of the track. For NAL unit structured video, the format of the list item is specified in ISO/IEC 14496-15.
[0180] The method and apparatus of an example embodiment may be utilized in a wide variety of systems, including systems that rely upon the compression and decompression of media data and possibly also the associated metadata. In at least an embodiment, however, the method and apparatus are configured to train or finetune a decoder-side neural network. In this regard, FIG. 6 depicts an example of such a system 600 that includes a source 602 of media data and associated metadata. The source 602 may be, in an embodiment, a server. However, the source may be embodied in other manners when desired. The source 602 is configured to stream the media data and associated metadata to a client device 604. The client device may be embodied by a media player, a multimedia system, a video system, a smart phone, a mobile telephone or other user equipment, a personal computer, a tablet computer or any other computing device configured to receive and decompress the media data and process associated metadata. In the illustrated embodiment, media data and metadata are streamed via a network 606, such as any of a wide variety of types of wireless networks and/or wireline networks. The client device is configured to receive structured information including media, metadata and any other relevant representation of information including the media and the metadata and to decompress the media data and process the associated metadata (e.g. for proper playback timing of decompressed media data).
[0181] An apparatus 700 is provided in accordance with an example embodiment as shown in FIG. 7. In an embodiment, the apparatus of FIG. 7 may be embodied by the source 602, such as a file writer which, in turn, may be embodied by a server, that is configured to stream a compressed representation of the media data and associated metadata. In an alternative embodiment, the apparatus may be embodied by the client device 604, such as a file reader which may be embodied, for example, by any of the various computing devices described above. In either of these embodiments and as shown in FIG. 7, the apparatus of an example embodiment includes, is associated with or is in communication with a processing circuitry 702, one or more memory devices 704, a communication interface 706 and optionally a user interface. [0182] The processing circuitry 702 may be in communication with the memory device 704 via a bus for passing information among components of the apparatus 700. The memory device may be non- transitory and may include, for example, one or more volatile and/or non-volatile memories. In other words, for example, the memory device may be an electronic storage device (e.g., a computer readable storage medium) comprising gates configured to store data (e.g., bits) that may be retrievable by a machine (e.g., a computing device like the processing circuitry). The memory device may be configured to store information, data, content, applications, instructions, or the like for enabling the apparatus to carry out various functions in accordance with an example embodiment of the present disclosure. For example, the memory device could be configured to buffer input data for processing by the processing circuitry. Additionally or alternatively, the memory device could be configured to store instructions for execution by the processing circuitry.
[0183] The apparatus 700 may, in some embodiments, be embodied in various computing devices as described above. However, in some embodiments, the apparatus may be embodied as a chip or chip set. In other words, the apparatus may comprise one or more physical packages (e.g., chips) including materials, components and/or wires on a structural assembly (e.g., a baseboard). The structural assembly may provide physical strength, conservation of size, and/or limitation of electrical interaction for component circuitry included thereon. The apparatus may therefore, in some cases, be configured to implement an embodiment of the present disclosure on a single chip or as a single ‘system on a chip.’ As such, in some cases, a chip or chipset may constitute means for performing one or more operations for providing the functionalities described herein.
[0184] The processing circuitry 702 may be embodied in a number of different ways. For example, the processing circuitry may be embodied as one or more of various hardware processing means such as a coprocessor, a microprocessor, a controller, a digital signal processor (DSP), a processing element with or without an accompanying DSP, or various other circuitry including integrated circuits such as, for example, an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), a microcontroller unit (MCU), a hardware accelerator, a special-purpose computer chip, or the like. As such, in some embodiments, the processing circuitry may include one or more processing cores configured to perform independently. A multi-core processing circuitry may enable multiprocessing within a single physical package. Additionally or alternatively, the processing circuitry may include one or more processors configured in tandem via the bus to enable independent execution of instructions, pipelining and/or multithreading.
[0185] In an example embodiment, the processing circuitry 702 may be configured to execute instructions stored in the memory device 704 or otherwise accessible to the processing circuitry. Alternatively or additionally, the processing circuitry may be configured to execute hard coded functionality. As such, whether configured by hardware or software methods, or by a combination thereof, the processing circuitry may represent an entity (e.g., physically embodied in circuitry) capable of performing operations according to an embodiment of the present disclosure while configured accordingly. Thus, for example, when the processing circuitry is embodied as an ASIC, FPGA or the like, the processing circuitry may be specifically configured hardware for conducting the operations described herein. Alternatively, as another example, when the processing circuitry is embodied as an executor of instructions, the instructions may specifically configure the processing circuitry to perform the algorithms and/or operations described herein when the instructions are executed. However, in some cases, the processing circuitry may be a processor of a specific device (e.g., an image or video processing system) configured to employ an embodiment of the present invention by further configuration of the processing circuitry by instructions for performing the algorithms and/or operations described herein. The processing circuitry may include, among other things, a clock, an arithmetic logic unit (ALU) and logic gates configured to support operation of the processing circuitry.
[0186] The communication interface 706 may be any means such as a device or circuitry embodied in either hardware or a combination of hardware and software that is configured to receive and/or transmit data, including video bitstreams. In this regard, the communication interface may include, for example, an antenna (or multiple antennas) and supporting hardware and/or software for enabling communications with a wireless communication network. Additionally or alternatively, the communication interface may include the circuitry for interacting with the antenna(s) to cause transmission of signals via the antenna(s) or to handle receipt of signals received via the antenna(s). In some environments, the communication interface may alternatively or also support wired communication. As such, for example, the communication interface may include a communication modem and/or other hardware/software for supporting communication via cable, digital subscriber line (DSL), universal serial bus (USB) or other mechanisms.
[0187] In some embodiments, the apparatus 700 may optionally include a user interface that may, in turn, be in communication with the processing circuitry 702 to provide output to a user, such as by outputting an encoded video bitstream and, in some embodiments, to receive an indication of a user input. As such, the user interface may include a display and, in some embodiments, may also include a keyboard, a mouse, a joystick, a touch screen, touch areas, soft keys, a microphone, a speaker, or other input/output mechanisms. Alternatively or additionally, the processing circuitry may comprise user interface circuitry configured to control at least some functions of one or more user interface elements such as a display and, in some embodiments, a speaker, ringer, microphone and/or the like. The processing circuitry and/or user interface circuitry comprising the processing circuitry may be configured to control one or more functions of one or more user interface elements through computer program instructions (e.g., software and/or firmware) stored on a memory accessible to the processing circuitry (e.g., memory device, and/or the like).
[0188] SEI prefix indication SEI message
[0189] The SEI prefix indication SEI message has been specified, for example, in HEVC and VVC. The SEI prefix indication SEI message carries one or more SEI prefix indications for SEI messages of a particular value of SEI payload type (payloadType). Each SEI prefix indication is a bit string that follows the SEI payload syntax of that value of payloadType and includes a number of complete syntax elements starting from the first syntax element in the SEI payload.
[0190] Each SEI prefix indication for an SEI message of a particular value of payloadType indicates that one or more SEI messages of this value of payloadType are expected or likely to be present in the coded video sequence (CVS), and to start with the provided bit string. A starting bit string would typically include a true subset of an SEI payload of the type of SEI message indicated by the payloadType, and may include a complete SEI payload.
[0191] SEI prefix indications should provide sufficient information for indicating what type of processing is needed or what type of content is included. The former (type of processing) indicates decoder-side processing capabilities, e.g., whether some type of frame unpacking is needed. The latter (type of content) indicates, for example, whether the bitstream includes subtitle captions in a particular language.
[0192] SEI processing order SEI message
[0193] The SEI processing order SEI message has been described, for example, in document JVET-AA2027. The SEI processing order SEI message carries information indicating a preferred processing order, as determined by the encoder (e.g., the content producer), for different types of SEI messages that may be present in the bitstream. When an SEI processing order SEI message is present, it is present in the first access unit of the coded video sequence. The SEI processing order SEI message persists in decoding order from the current access unit until the end of the CVS. The SEI processing order SEI message comprises a list of pairs, each pair comprising a SEI payload type value po_sei_payload_type[ i ] and a processing order value po_sei_processing_order[ i ]. po_sei_payload_type[ i ] specifies a value of payloadType for the i-th SEI message for which information is provided in the SEI processing order SEI message. po_sei_processing_order[ i ] indicates the preferred order of processing any SEI message with payloadType equal to po_sei_payload_type[ i ]. po_sei_processing_order[ m ] greater than 0 and less than po_sei_processing_order[ n ] indicates any SEI message with payloadType equal to po_sei_payload_type[ m ], when present, should be processed before any SEI message with payloadType equal to po_sei_payload_type[ n ]. po_sei_processing_order[ m ] greater than 0 and equal to po_sei_processing_order[ n ] indicates that the preferred order of processing of SEI messages with payloadTypes equal to po_sei_payload_type[ m ] and po_sei_payload_type[ n ] is unknown, unspecified, or determined by external means. po_sei_processing_order[ i ] equal to 0 specifies that the preferred order of processing SEI messages with payloadType equal to po_sei_payload_type[ i ] is unknown, unspecified, determined by external means.
[0194] Neural-network post-filter characteristics (NNPFC) and neural- network post-filter activation (NNPFA) SEI messages
[0195] The neural-network post-filter characteristics (NNPFC) SEI message and the neural- network post-filter activation (NNPFA) SEI message have been described, for example, in document JVET-AA2006.
[0196] An NNPFC SEI message identifies an applicable post-processing filter associated with the nnpfc_id value. The use of applicable post-processing filters with different values of nnpfc_id for specific pictures is indicated with neural-network post-filter activation (NNPFA) SEI messages.
[0197] An NNPFC SEI message either specifies a base post-processing filter or includes a neural network update. A base post-processing filter is identified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within a CLVS. When there is no subsequent NNPFC SEI message that has the same nnpfc_id value as the base post-processing filter, the applicable postprocessing filter is the same as the base post-processing filter. Otherwise, the applicable post-processing filter is obtained by applying the update provided as an ISO/IEC 15938-17 bitstream in a subsequent NNPFC SEI message on top of the base post-processing filter.
[0198] The NNPFC SEI message comprises the nnpfc_id syntax element, which includes an identifying number that may be used to identify a post-processing filter. A base post-processing filter is the filter that is included in or identified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within a coded layer video sequence (CLVS). When there is another NNPFC SEI message that has the same nnpfc_id value and different content than the NNPFC SEI message that defines the base post-processing filter, the base post-processing filter is updated by decoding the coded neural network bitstream in that NNPFC SEI message to obtain a post-processing filter associated with the nnpfc_id value. Otherwise, the post-processing processing filter associated with the nnpfc_id value is assigned to be the same as the base post-processing filter.
[0199] The NNPFC SEI message comprises nnpfc_mode_idc syntax element, the semantics of which may be defined as follows: nnpfc_mode_idc equal to 0 specifies that the base post-processing filter associated with the nnpfc_id value is determined by external means, e.g., means specified outside of the NNPFC SEI message. nnpfc_mode_idc equal to 1 indicates that this SEI message includes an ISO/IEC 15938-17 bitstream that specifies the base post-processing filter or updates the base post-processing filter with the same nnpfc_id value. The ISO/IEC 15938-17 bitstream may be at the end of the NNPFC SEI message. In other words, no syntax elements may follow the ISO/IEC 15938-17 bitstream within the NNPFC SEI message. nnpfc_mode_idc equal to 2 specifies that the base post-processing filter associated with the nnpfc_id value is a neural network identified by the uniform resource identifier (URI) nnpfc_uri with the format identified by the tag URI nnpfc_tag_uri.
[0200] The NNPFC SEI message may also comprise:
Purpose of the post- processing filter, such as: o Visual quality improvement. o Chroma upsampling from the 4:2:0 chroma format to the 4:2:2 or 4:4:4 chroma format, or from the 4:2:2 chroma format to the 4:4:4 chroma format. o Increasing the width or height of the cropped decoded output picture without changing the chroma format o Increasing the width or height of the cropped decoded output picture and upsampling the chroma format..
Formatting of the input tensors that are given as input to the neural network inference. Formatting of the output tensors that are resulting from the neural network inference. Characterization of the complexity of the neural network.
[0201] The NNPFA SEI message specifies the neural-network post-processing filter that may be used for post-processing filtering for the current picture. The NNPFA SEI message comprises the nnpfa_id syntax element, which specifies that the neural-network post-processing filter with nnpfc_id equal to nnfpa_id may be used for post-processing filtering for the current picture. [0202] ISO base media file format
[0203] Some concepts, structures, and specifications of ISOBMFF are described below as an example of a container file format, based on which some embodiments may be implemented. The features of the disclosure are not limited to ISOBMFF, but rather the description is given for one possible basis on top of which at least some embodiments may be partly or fully realized.
[0204] A basic building block in the ISO base media file format is called a box. Each box has a header and a payload. The box header indicates the type of the box and the size of the box in terms of bytes. A box may enclose other boxes, and the ISO file format specifies which box types are allowed within a box of a certain type. Furthermore, the presence of some boxes may be mandatory in each file, while the presence of other boxes may be optional. Additionally, for some box types, it may be allowable to have more than one box present in a file. Thus, the ISO base media file format may be considered to specify a hierarchical structure of boxes.
[0205] According to the ISO family of file formats, a file includes media data and metadata that are encapsulated into boxes. Each box is identified by a four character code (4CC) and starts with a header which informs about the type and size of the box.
[0206] In files conforming to the ISO base media file format, the media data may be provided in a media data box ('mdat', also called MediaDataBox) and the movie box ('moov', also called MovieBox) may be used to enclose the metadata. In some examples, for a file to be operable, both of the ‘mdat’ and ‘moov’ boxes may be required to be present. The movie box may include one or more tracks, and each track may reside in one corresponding track box ('trak', may also be called TrackBox). A track may be one of the many types, including a media track that refers to samples formatted according to a media compression format (and its encapsulation to the ISO base media file format).
[0207] The 'trak' box includes a Sample Table box. The Sample Table box includes, for example, time and data indexing of the media samples in a track. The Sample Table box is required to include a Sample Description box. The Sample Description box includes an entry count field, specifying the number of sample entries included in the box. The Sample Description box is required to include at least one sample entry. The sample entry format depends on the handler type for the track. Sample entries give detailed information about the coding type used and any initialization information needed for that coding. [0208] Movie fragments may be used, for example, when recording content to ISO files, for example, in order to avoid losing data when a recording application crashes, runs out of memory space, or some other incident occurs. Without movie fragments, data loss may occur because the file format may require that all metadata, for example, the movie box, be written in one contiguous area of the file. Furthermore, when recording a file, there may not be sufficient amount of memory space (e.g., random access memory RAM) to buffer a movie box for the size of the storage available, and re-computing the contents of a movie box when the movie is closed may be too slow. Moreover, movie fragments may enable simultaneous recording and playback of a file using a regular ISO file parser. Furthermore, a smaller duration of initial buffering may be required for progressive downloading, for example, simultaneous reception and playback of a file when movie fragments are used, and the initial movie box is smaller compared to a file with the same media content but structured without movie fragments.
[0209] A movie fragment feature may enable splitting the metadata that otherwise might reside in the movie box into multiple pieces. Each piece may correspond to a certain period of time of a track. In other words, the movie fragment feature may enable interleaving file metadata and media data. Consequently, the size of the movie box may be limited, and the use cases mentioned above be realized.
[0210] A MovieBox may include a MovieExtendsBox ('mvex'). When present, presence of the MovieExtendsBox warns readers that there might be movie fragments in this file or stream. To know of all samples in the tracks, movie fragments are obtained and scanned in order, and their information logically added to information in the MovieBox. A MovieExtendsBox includes one TrackExtendsBox per track. A TrackExtendsBox includes default values used by the movie fragments. Some examples of the default values that can be given in TrackExtendsBox, include but are not limited to: default sample description index (e.g., default sample entry index), default sample duration, default sample size, and default sample flags. Sample flags include dependency information, such as when the sample depends on other sample(s), when other sample(s) depend on the sample, and when the sample is a sync sample.
[0211] In some examples, the media samples for the movie fragments may reside in an mdat box, when the movie fragments are in the same file as the moov box. For the metadata of the movie fragments, however, a moof box (also called MovieFragmentBox) may be provided. The moof box may include information for a certain duration of playback time that would previously have been in the moov box. The moov box may still represent a valid movie on its own, but in addition, it may include an mvex box indicating that movie fragments will follow in the same file. The movie fragments may extend the presentation that is associated to the moov box in time. [0212] Within the movie fragment there may be a set of track fragments, including anywhere from zero to a plurality per track. The track fragments may in turn include anywhere from zero to a plurality of track runs, each of which document is a contiguous run of samples for that track. Within these structures, many fields are optional and may have default values. The metadata that may be included in the moof box may be limited to a subset of the metadata that may be included in a moov box and may be coded differently in some cases. Details regarding the boxes that can be included in a moof box may be found from the ISO base media file format specification.
[0213] The track reference mechanism may be used to associate tracks with each other. The TrackReferenceBox includes box(es), each of which provides a reference from the including track to a set of other tracks. These references are labeled through the box type (e.g., the four-character code of the box) of the included box(es).
[0214] In ISOBMFF, a track group enables grouping of tracks based on certain characteristics or the tracks within a group have a particular relationship. Track grouping, however, does not allow any image items in the group. A track group box (also known as TrackGroupBox) may be present in a TrackBox and may include boxes that are derived from TrackGroupTypeBox, which is a box whose box payload starts with track_group_id and whose box type (also referred to as track_group_type) defines the track group type.
[0215] The pair of track_group_id and track_group_type identifies a track group within a file. The tracks that include a particular TrackGroupTypeBox having the same value of track_group_id and track_group_type belong to the same track group.
[0216] The TrackGroupDescriptionBox may be included in the MovieBox. The TrackGroupDescriptionBox provides an array of TrackGroupEntryBoxes, where each TrackGroupEntryBox provides detailed characteristics of a particular track group. The syntax of the TrackGroupEntryBox is determined by track_group_entry_type. TrackGroupEntryBox is mapped to the track group by a unique track_group_entry_type that is associated with a track_group_type. More than one TrackGroupEntryBox with the same track_group_entry_type and different track_group_id may be present in TrackGroupDescriptionBox.
[0217] Files conforming to the ISOBMFF may include any non-timed objects, referred to as items, meta items, or metadata items, in a meta box (four-character code: ‘meta’), also referred to as MetaBox or metabox. While the name of the meta box refers to metadata, items can generally include metadata or media data. The meta box may reside at the top level of the file, within a movie box (four- character code: ‘moov’), and within a track box (four-character code: ‘trak’), but at most one meta box may occur at each of the file level, movie level, or track level.
[0218] The meta box may be required to include a ‘hdlr’ box indicating the structure or format of the ‘meta’ box contents.
[0219] The meta box may list and characterize any number of items that can be referred and each one of them can be associated with a file name and are uniquely identified with the file by item identifier (item_id) which is an integer value. The metadata items may be for example stored in the ItemDataBox fidaf box) of the meta box or in an 'mdat' box or reside in a separate file. If the metadata is located external to the file, then its location may be declared by the DatalnformationBox (four-character code: ‘dinf’).
[0220] In the specific case that the metadata is formatted using extensible Markup Language (XML) syntax and is required to be stored directly in the MetaBox, the metadata may be encapsulated into either the XMLBox (four-character code: ‘xml ‘) or the BinaryXMLBox (four-character code: ‘bxml’).
[0221] An item may be stored as a contiguous byte range, or it may be stored in several extents, each being a contiguous byte range. In other words, items may be stored fragmented into extents, e.g., to enable interleaving. An extent is a contiguous subset of the bytes of the resource. The resource can be formed by concatenating the extents.
[0222] The ItemPropertiesBox enables the association of any item with an ordered set of item properties. Item properties may be regarded as small data records. The ItemPropertiesBox include two parts: ItemPropertyContainerBox that includes an implicitly indexed list of item properties, and one or more ItemProperty AssociationBox(es) that associate items with item properties.
[0223] Item references may be indicated in the ItemReferenceBox (’iref). The ItemReferenceBox allows the linking of one item to others via typed references. All the references for one item of a specific type are collected into a SingleltemTypeReferenceBox, whose type (in the box header) is the reference type, and which has a from_item_ID field indicating which item is linked. The items linked to are then represented by an array of to_item_IDs. All these single item type reference boxes are then collected into the ItemReferenceBox. [0224] HandlerBox, when present within a MetaBox, declares the structure or format of the MetaBox contents. The MetaBox may also be referred to as ‘meta box’ in some embodiments. There is a general handler for metadata streams of any type; the specific format is identified by the sample entry, as for video or audio, for example.
[0225] HandlerProperty provides a mapping of a media handler with an item in a MetaBox. When a HandlerBox is present, it applies to all items without a HandlerProperty and may provide additional requirements on items with a HandlerProperty with different handler lype than the one in the HandlerBox.
[0226] DataReferenceBox ('dref box) includes a list of boxes that declare the potential location(s) of the media data referred to by the file. DataReferenceBox is included by DatalnformationBox, which in turn is included by MedialnformationBox or MetaBox. When included in the MedialnformationBox, each sample entry of the track includes a data reference index referring to a list entry of the list of box(es) in the DataReferenceBox. When included in the MetaBox, the ItemLocationBox gives, for each item, the data reference index referring to a list entry of the list of box(es) in the DataReferenceBox. The box(es) in the DataReferenceBox are extended from FullBox, e.g., included the version and the flags field in the box header. As an example, DataReferenceBox may comprise DataEntryUrlBox and DataEntryUrnBox, which provide a uniform resource locator (URL) and a uniform resource name (URN) data reference, respectively. When the least significant bit of the flags field of either DataEntryUrlBox or DataEntryUrnBox is equal 1, the respective data reference refers to the including file itself and no URL or URN string is provided within the DataEntryUrlBox or the DataEntryUrnBox.
[0227] The Entity grouping is similar to track grouping but enables grouping of both tracks and items in the same group. The entities in an entity group share a particular characteristic or have a particular relationship, as indicated by the grouping type.
[0228] Entity groups are indicated in GroupsListBox. Entity groups specified in GroupsListBox of a file-level MetaBox refer to tracks or file-level items. Entity groups specified in GroupsListBox of a movie-level MetaBox refer to movie-level items. Entity groups specified in GroupsListBox of a tracklevel MetaBox refer to track-level items of that track.
[0229] GroupsListBox includes EntityToGroupBoxes, each specifying one entity group.
[0230] The syntax of EntityToGroupBox in ISOBMFF is defined as follows: aligned(8) class EntityToGroupBox(grouping_type, version, flags) extends FullBox(grouping_type, version, flags) { unsigned int(32) group_id; unsigned int(32) num_entities_in_group; for(i=0; i<num_entities_in_group; i++) unsigned int(32) entity _id;
}
[0231] group_id is a non-negative integer assigned to a particular grouping that may not be equal to any group_id value of any other EntityToGroupBox, any item_ID value of the hierarchy level (e.g., a file, a movie or a track) that includes the GroupsListBox, or any track_ID value (when the GroupsListBox is included in the file level).
[0232] num_entities_in_group specifies a number of entity _id values mapped to this entity group.
[0233] entity _id is resolved to an item, when an item with item_ID equal to entity _id is present in the hierarchy level (e.g., a file, a, movie or a track) that includes the GroupsListBox, or to a track, when a track with track_ID equal to entity _id is present and the GroupsListBox is included in the file level.
[0234] Restricted media tracks in ISO base media file format
[0235] Restricted media sample entries, such as the restricted video ('resv') sample entry, and the restricted media mechanism have been specified for the ISOBMFF in order to handle situations where the file author requires certain actions on the player or Tenderer before or after decoding of a media track. Players not recognizing or not capable of processing the required actions are stopped from decoding or rendering the restricted video tracks. The 'resv' sample entry mechanism applies to any type of video codec. A RestrictedSchemelnfoBox is present in the sample entry of 'resv' tracks and comprises an OriginalFormatBox, SchemeTypeBox, and SchemelnformationBox. The original sample entry type that would have been unless the 'resv' sample entry type were used is included in the OriginalFormatBox. The SchemeTypeBox provides an indication which type of processing is required in the player to process the video. The SchemelnformationBox comprises further information of the required processing. The scheme type may impose requirements on the contents of the SchemelnformationBox. For example, the stereo video scheme indicated in the SchemeTypeBox indicates that when decoded frames either include a representation of two spatially packed constituent frames that form a stereo pair (frame packing) or only one view of a stereo pair (left and right views in different tracks). StereoVideoBox may be included in a SchemelnformationBox to provide further information, e.g., on which type of frame packing arrangement has been used (e.g., side-by-side or topbottom).
[0236] ISO/IEC 14496-15 specifies a restricted media scheme for SEI-message-based postprocessing, labelled as 'aSEI'. For the case of signalling of SEI-driven post-processing, a file author may list occurring SEI message IDs and classify them into two categories: those that are deemed required by the file author for correct playback, and others. The occurrence of either type of SEI messages may be signalled in the SEI information box.
[0237] The SEI information box (a.k.a. SeilnformationBox) documents the SEI payload IDs in a stream. It may be included in a VisualSampleEntry or in a SchemelnformationBox.
[0238] Sample group in ISO base media file format
[0239] A sample grouping in the ISO base media file format and its derivatives may be defined as an assignment of each sample in a track to be a member of one sample group, based on a grouping criterion. A sample group in a sample grouping is not limited to being contiguous samples and may include non-adjacent samples. As there may be more than one sample grouping for the samples in a track, each sample grouping may have a type field grouping type to indicate the type of grouping. Sample groupings may be represented by two linked data structures: (1) a SampleToGroupBox ('sbgp' box) represents the assignment of samples to sample groups; and (2) a SampleGroupDescriptionBox fsgpd' box) includes sample group (description) entries for describing the properties of samples mapped to this entry. There may be multiple instances of the SampleToGroupBox and SampleGroupDescriptionBox based on different grouping criteria. These may be distinguished by a type field used to indicate the type of grouping. SampleToGroupBox may additionally comprise a grouping_type_parameter field that can be used e.g., to indicate a sub-type of the grouping.
[0240] An essential sample group description is a sample group description for which the version field is equal to 3, and the associated sample group is also referred to as an essential sample group. An essential sample group description describes essential information for the associated samples, and parsers are not allowed to attempt to process any track for which unrecognized sample group descriptions marked as essential are present. [0241] The essential descriptions hierarchy sample group fesgh') indicates the processing order of the essential sample group descriptions applying to a given sample. This sample group description is an essential sample group description and uses version 3 of the SampleGroupDescriptionBox.
[0242] Each essential sample group description, except the essential descriptions hierarchy sample group itself, is listed in the EssentialDescriptionsHierarchyEntry.
[0243] Samples associated with essential sample groups use a restricted sample entry indicating the original media type (e.g. 'resv', 'resa') with a scheme lype equal to 'essg'. In a sample entry, there is at most one sample entry transformation with a scheme lype equal to 'essg'. When such a transformation is present, the following applies:
The transformation is the first sample entry transformation.
There is no other sample entry transformations than protection and restricted media transformations.
There is at most one sample entry transformation of type protection.
[0244] The transformations given in sample_group_description_type are listed in the order in which a file reader applies each transformation: any sample processing described by a sample group of type sample_group_description_type[i] is applied before any sample processing described by a sample group of type sample_group_description_type[i+l].
[0245] In the sample_group_description_type list, the following transformation values are allowed:
'stsd' : indicates the position of the decoding process in the transformation chain.
'cenc' : indicates the position of the decryption in the transformation chain.
Any scheme lype value included in the SchemeTypeBox of the RestrictedSchemelnfoBox of the track that includes this sample group; this indicates the position of the respective restricted media transformation in the transformation chain.
[0246] When 'stsd' is absent from the list of sample_group_description_type, all listed transformations apply to decoded samples. When 'cenc' is present in the list, 'stsd' is also present.
[0247] The codecs MIME parameter may include identification of essential sample groups as follows: When a restricted media track is implied by a sample entry and the transformation type indicates an essential sample group (scheme_type equal to 'essg'), the value of the codecs MIME parameter is appended by the four-character code 'essg' followed by a star ('*'), further followed by the four-character codes listed in the essential descriptions hierarchy sample description, from the first entry up to but excluding the first occurrence of 'stsd' or 'cenc'. A star ('*') is used to separate the four-character codes listed in the essential descriptions hierarchy sample description.
[0248] For files including essential sample group descriptions, the 'essential' MIME parameter, when used, is composed of one or more comma-separated essential hierarchy descriptions. Each essential hierarchy description is composed of one or more four-character code of essential sample group descriptions, separated with a dot. When the ‘codecs’ parameter includes description of the transformation used, the listed four-character codes are the ones listed in the essential descriptions hierarchy sample group description, in the same order, from the first code following the last occurrence of 'stsd' until the last listed code. Otherwise, the listed four-character codes are the ones listed in the essential descriptions hierarchy sample group description in the same order.
[0249] Dynamic adaptive streaming over HTTP (DASH)
[0250] Recently, Hypertext Transfer Protocol (HTTP) has been widely used for the delivery of real-time multimedia content over the Internet, such as in video streaming applications. Several commercial solutions for adaptive streaming over HTTP, such as Microsoft® Smooth Streaming, Apple® Adaptive HTTP Live Streaming and Adobe® Dynamic Streaming, have been launched as well as standardization projects have been carried out. Adaptive HTTP streaming (AHS) was first standardized in Release 9 of 3rd Generation Partnership Project (3GPP) packet-switched streaming (PSS) service (3GPP TS 26.234 Release 9: “Transparent end-to-end packet-switched streaming service (PSS); protocols and codecs”). MPEG took 3GPP AHS Release 9 as a starting point for the MPEG DASH standard (ISO/IEC 23009-1: “Dynamic adaptive streaming over HTTP (DASH)-Part 1: Media presentation description and segment formats,” International Standard, 2nd Edition, , 2014). 3GPP continued to work on adaptive HTTP streaming in communication with MPEG and published 3GP- DASH (Dynamic Adaptive Streaming over HTTP; 3GPP TS 26.247: “Transparent end-to-end packet- switched streaming Service (PSS); Progressive download and dynamic adaptive Streaming over HTTP (3GP-DASH)”. MPEG DASH and 3GP-DASH are technically close to each other and may therefore be collectively referred to as DASH. Some concepts, formats, and operations of DASH are described below as an example of a video streaming system, wherein the embodiments may be implemented. The embodiments of the invention are not limited to DASH, but rather the description is given for one possible basis on top of which the invention may be partly or fully realized.
[0251] In DASH, the multimedia content may be stored on an HTTP server and may be delivered using HTTP. The content may be stored on the server in two parts: Media Presentation Description (MPD), which describes a manifest of the available content, its various alternatives, their URL addresses, and other characteristics; and segments, which include the actual multimedia bitstreams in the form of chunks, in a single file or multiple files. The MDP provides the necessary information for clients to establish a dynamic adaptive streaming over HTTP. The MPD includes information describing media presentation, such as an HTTP- uniform resource locator (URL) of each Segment to make GET Segment request. To play the content, the DASH client may obtain the MPD e.g. by using HTTP, email, thumb drive, broadcast, or other transport methods. By parsing the MPD, the DASH client may become aware of the program timing, media-content availability, media types, resolutions, minimum and maximum bandwidths, and the existence of various encoded alternatives of multimedia components, accessibility features and required digital rights management (DRM), media-component locations on the network, and other content characteristics. Using this information, the DASH client may select the appropriate encoded alternative and start streaming the content by fetching the segments using e.g. HTTP GET requests. After appropriate buffering to allow for network throughput variations, the client may continue fetching the subsequent segments and also monitor the network bandwidth fluctuations. The client may decide how to adapt to the available bandwidth by fetching segments of different alternatives (with lower or higher bitrates) to maintain an adequate buffer.
[0252] In DASH, hierarchical data model is used to structure media presentation as follows. A media presentation includes a sequence of one or more Periods, each Period includes one or more Groups, each Group includes one or more adaptation sets, each adaptation sets includes one or more representations, each representation includes one or more segments. A representation is one of the alternative choices of the media content or a subset thereof typically differing by the encoding choice, e.g. by bitrate, resolution, language, codec, etc. The segment includes certain duration of media data, and metadata to decode and present the included media content. A segment is identified by a URI and can typically be requested by a HTTP GET request. A segment may be defined as a unit of data associated with an HTTP-URL and optionally a byte range that are specified by an MPD.
[0253] The DASH MPD complies with Extensible Markup Language (XML) and is therefore specified through elements and attributes as defined in XML.
[0254] Preselection in ISOBMFF and MPEG-DASH
[0255] Features as described herein may relate to preselection in ISOBMFF. Preselection may be defined as set of one or more tracks representing one version of the media presentation for simultaneous decoding or presentation. [0256] ISOBMFF defines a preselection group box. A track grouping of type 'pres' may indicate that a track contributes to a preselection. The pair of track_group_id and track_group_type identifies a track group. The tracks that include a particular TrackGroupTypeBox having the same value of track_group_id and track_group_type 'pres' belong to the same preselection group.
[0257] The presence of a TrackGroupTypeBox with track_group_type equal to 'pres' (which may also be referred to as a PreselectionGroupBox) in a track indicates that this track contributes to a preselection. All the tracks that have a track group with track_group_type equal to 'pres' and a particular value of track_group_id are part of the same preselection. The particular value of track_group_id may also be referred to as the ID of the preselection. This means that a preselection is uniquely identified by the track_group_id of the track group. When multiple tracks contribute to a preselection, the optionally present PreselectionProcessingBox may provide information on how to process the track including this box in the context of the preselection and relative to other tracks. Consequently, the content of the PreselectionProcessingBox may differ for each track within a preselection. Preselections made from only one track may not require any track-related processing. In this case, the PreselectionProcessingBox is typically not present in the PreselectionGroupBox.
[0258] The syntax of PreselectionGroupBox is defined below: aligned(8) class PreselectionGroupBox extends TrackGroupTypeBox('pres')
{
PreselectionProcessingBox preselection_processing; // optional
}
[0259] preselection_processing is an instance of the PreselectionProcessingBox, providing information needed for processing the including track in the context of the preselection.
[0260] The PreselectionProcessingBox box may include information about how the tracks contributing to the preselection may be processed. Media type specific boxes may be used to describe further processing.
[0261] The syntax of PreselectionProcessingBox is defined below: aligned(8) class PreselectionProcessingBox extends FullBoxfprsp', version=0, flags ){ unsigned int(8) track_order; unsigned int(l) sample_merge_flag; unsigned int(7) reserved; 11 further attributes and Boxes defining additional processing of
// the track contributing to the preselection }
[0262] The corresponding semantics of the syntax elements may be defined as follows. track_order may define the order of this track relative to other tracks in the preselection, as described below. sample_merge_flag equal to 1 may indicate that this track is enabled to be merged with another track, as described below. Sample entry specific specifications may require that the tracks for a preselection be provided to the respective decoder instances in a specific order. Since other means, such as the track_id, are not reliable for this purpose, the track_order may be used to order tracks in a preselection relative to each other. A lower number may indicate that at a given time, the sample of the including track is provided to the decoder before the sample with the same given of other tracks with higher number for track_order. If two tracks in a preselection have their track_order set to the same value, or if the preselection processing box is absent for at least one of the tracks, the order of these tracks may not be relevant for the preselection, and samples may be provided to the decoder in any order.
[0263] A merge group may be defined as a group of tracks, sorted according to track_order, where one track with the sample_merge_flag set to 0 is followed by a group of consecutive tracks with the sample_merge_flag set to 1. All tracks of a merge group may be of the same media type, and all samples may be time-aligned.
[0264] When the sample entry type is associated with a codec-specific process to merge samples of a preselection, this process may be used. When the tracks in the merge group are all of sample entry type of “mhm2” (MPEG-H 3D Audio), the merging process is defined in ISO/IEC 23008-3:2019. Tracks in a merge group may have different sample entry types. When the sample entry type is not associated with a codec-specific process to merge samples of a preselection, the following process may be used. Merging within the merge group may comprise forming tuples of track samples with the same time stamp across contributing tracks. The ordering of samples within the tuple may be determined by track_order. These tuples may be formed by byte-wise concatenation of the samples, resulting in a single sample having the respective time stamp assigned. When generation of new tracks is targeted, each merge group may result in a separate output track conformant to a media type derived from the media types of the merged tracks. For tracks not part of a merge group.
[0265] Preselections may be qualified, for example, by language, kind, or media specific attributes, like audio rendering indication(s), audio interactivity, and/or channel layout(s). Attributes signaled in a PreselectionTrackGroupEntryBox may take precedence over attributes signaled in contributing tracks.
[0266] PreselectionTrackGroupEntryBox may describe only track groups identified by track_group_type equal to 'prse'. All preselections with at least one contributing track having the track_in_movie flag set to 1 may be qualified by PreselectionTrackGroupEntryBoxes. Otherwise, the presence of the PreselectionTrackGroupEntryBoxes may be optional. All attributes uniquely qualifying a preselection may be present in the PreselectionTrackGroupEntryBox of the preselection.
[0267] An example of a PreselectionTrackGroupEntryBox may be as follows: aligned(8) class PreselectionTrackGroupEntryBox extends TrackGroupEntryBox('prse', version, flags)
{ unsigned int(8) numTracks; utf8string preselection_tag; if (flags & 1) { unsigned int(8) selection_priority;
} if (flags & 2) { unsigned int(8) segment_order;
}
// Boxes describing the preselection
}
[0268] This box includes information on what experience is available when this preselection is selected.
[0269] Boxes suitable to describe a preselection include, but are not limited to, the following list of boxes: the audio element box; the audio element selection box; the extended language tag; the user data box; the track kind; the label box; the audio rendering indication; and/or the channel layout. If a UserDataBox is included in a PreselectionTrackGroupEntryBox, then it may not carry any of the above boxes.
[0270] numTracks may specify the number of non-alternative tracks grouped by a preselection track group. A track grouped by this preselection track group may be a track that has the 'pres' track group with track_group_id equal to the ID of this preselection. The number of non-alternative tracks grouped by this preselection track group may be the sum of the following: the number of tracks that have alternate_group equal to 0 and are grouped by this preselection track group; and/or the number of unique non-zero alternate_group values in all tracks that are grouped by this preselection track group. The value of numTracks may be greater than or equal to the number of non-alternative tracks grouped by this preselection track group in this file. A value equal to 0 may indicate that the number of tracks grouped by this track group is unknown or not essential for processing the track group.
[0271] It may be noted that the value of numTracks may be greater than the number of tracks containing a PreselectionGroupBox with the same track_group_id in this file when the preselection is split into multiple files.
[0272] It may be noted that when a player has access to fewer non-alternative tracks grouped by this preselection track group than indicated by numTracks, the player might need to omit the tracks grouped by this preselection track group.
[0273] preselection_tag may be a codec specific value that a playback system may provide to a decoder to uniquely identify one out of several preselections in the media.
[0274] selection_priority may be an integer that declares the priority of the preselection in cases where no other differentiation, such as through the media language, is possible. A lower number may indicate a higher priority.
[0275] segment_order may specify, when present, an order rule of segments that may be suggested to be followed for ordering received segments of the Preselection. The following values may be specified with semantics according to ISO/IEC 23009-1: 0: undefined; 1: time-ordered; 2: fully-ordered. Other values may be reserved. When segment_order is not present, its value may be inferred to be equal to 0.
[0276] It may be noted that not all tracks contributing to the playout of a preselection may be delivered in the same file.
[0277] It may be noted that the kind box might utilize the Role scheme defined in ISO/IEC 23009- 1, as it provides a commonly used scheme to describe characteristics of preselections. [0278] It may be noted that this box may carry information about the initial experience of the preselection in the referenced tracks. The preselection experience may change during the playback of these tracks (e.g. audio language may change during playback). These changes are not subject to the information presented in this box.
[0279] Further media type specific boxes may be used to describe properties of the preselection.
[0280] Features as described herein may relate to preselection in MPEG-DASH. The concept of preselections was added to MPEG-DASH in order to enable the combination of different adaptation sets into a single decoding instance and user experience.
[0281] A DASH preselection defines a subset of media components of an MPD that are expected to be consumed jointly by a single decoder instance, wherein consuming may comprise decoding and rendering. The adaptation set that includes the main media component for a preselection is referred to as main adaptation set. In addition, each preselection may include one or multiple partial adaptation sets. Partial adaptation sets may need to be processed in combination with the main adaptation set. A main adaptation set and partial adaptation sets may be indicated by one of the two means: a preselection descriptor or a preselection element.
[0282] Preselections may be used to select experiences. Each preselection may reference one or more media content components within one or multiple adaptation sets. Preselection may be defined as a set of media content components that are intended to be consumed jointly. A media content component may be a single continuous component of the media content with an assigned media content component type. The media content component type may be a single type of media content, for example audio, video, or text.
[0283] The concept of preselections was initially considered for the purpose of enabling Next Generation Audio (NGA) codecs to signal suitable combinations of audio elements that are offered in different Adaptation Sets. However, the preselection concept is introduced in a generic manner such that it can also be applicable to and used by other media types and codecs. Preselections define user experiences that may be selected by the DASH Client. Each Preselection may be uniquely identifiable and distinguishable, (e.g. by language). A preselection may encompass a subset of media components, such that the media components may be selected and combined into a complete experience. Preselections may be used to reference a set of Representations from multiple Adaptation Sets in order to produce a complete experience. Preselections may also be used to indicate a pre-defined experience at the elementary- stream level (e.g.,. the DASH Client may select a pre-defined experience and provide the selection to the media engine). Preselections may be uniquely identified by a preselection tag. Users/Codecs using this tag functionality may be encouraged to provide more information on how tags defined in the MPD map to functionality in the specific codec.
[0284] Preselections may have equivalent annotation parameters to adaptation sets and may always be assigned exactly one media type.
[0285] Media components may be mapped to adaptation sets in multiple ways: by a one-to-one mapping between media components and Adaptation Sets; by the inclusion of multiple media components in a single adaptation set, where all encoded versions of the media components may be multiplexed on the file-container level; and/or by the inclusion of multiple media components in a single adaptation set, where all encoded versions of the media components may be multiplexed on the elementary-stream level.
[0286] When the adaptation set includes a single media component, then the media component may be referenced by the @id of the adaptation set. When the adaptation set includes multiple media components multiplexed on the file- container level, then each media component may be mapped to a content component. For example, in the ISO BMFF case, a representation may include multiple tracks, and each track may be mapped to a content component. Therefore, media components may be referenced by the @id of an adaptation set, or the @ id of a content component. When preselections reference content components, the @id of adaptation sets and content components may be unique within the scope of a period.
[0287] When the adaptation set includes multiple media components multiplexed at the elementary-stream level, then a pre-defined experience may be referenced by the Preselection. For example, in the ISOBMFF case, a pepresentation may include a single track of multiple media components that may be referenced by the @id of the adaptation set. Multiple preselections may reference the adaptation set and select a pre-defined experience by passing the presentation tag to the media engine, along with the media stream.
[0288] Within a preselection, two types of adaptation sets may be differentiated, a main adaptation set and a partial adaptation set. The main adaptation set is a representation of this adaptation set that may be needed for playback of the preselection. In particular for ISO BMFF, the initialization segment of such a representation is needed for playback of the preselection. In other words, the main adaptation set is the adaptation set that includes the initialization segment for the complete experience. Each preselection may reference a main adaptation set and may reference zero, one, or more other adaptation sets. In the context of preselection, the term "main adaptation set" is not to be confused with an adaptation set that has assigned the main role. A partial adaption set is a representation of this adaptation set that may only be consumable together with the main adaptation set(s) within this preselection. Again, in particular for ISO BMFF, the initialization segment of a representation of the main adaptation set is needed for playback.
[0289] Preselections, main adaptation set and partial adaptation sets may be defined by one of the two means: a preselection element, or a preselection descriptor. A preselection descriptor may enable simple configurations and may preserve backward compatibility, but may not be suitable for advanced use cases. The preselection descriptor may be used for two purposes: to indicate that an adaptation Set is part of a preselection; or, optionally, to provide instructions on how to combine the adaptation set with other adaptation sets to form a preselection.
[0290] The preselection descriptor may be either an essential property descriptor or a supplemental property descriptor with a @schemeIdURI of "um:mpeg:dash:preselection:2016".
[0291] For simple use cases, the preselection descriptor may be used to provide instructions on how to combine the adaptation set with other adaptation sets to form a preselection. The descriptor may only be present at the adaptation set level.
[0292] The @ value attribute of the descriptor may provide two fields, separated by a comma: the preselection tag, and the id of the adaptation sets or content components referenced by this preselection, for example, as a white space separated list in processing order. The first id may reference the main Adaptation Set. The syntax for the value attribute of the preselection descriptor may follow the PRESELECTION-DESCRIPTOR-VALUE as defined in the following ABNF notation according to IETF RFC 5234:
PRESELECTION-DESCRIPTOR- VALUE = TAG- VALUE ID-LIST
ID-LIST = ID-VALUE [ WHITESPACE ID- VALUE]
TAG- VALUE = STRING
ID- VALUE = 1* DIGIT
STRING = *VCHAR
[0293] When the descriptor is present, but the value field is absent, the adaptation set may be referenced by at least one preselection. When the value field is present, then this descriptor may identify a preselection. The tag may be assigned, and the preselection may include the ids of the adaptation sets that are part of this preselection. In this case, the multi-segment track conformance rule may apply. When the descriptor is present and used with the essential descriptor, then this may indicate that the adaptation set is only consumable as part of a preselection. When the descriptor is present and used with the supplemental descriptor, then this may indicate that the adaptation set is also consumable independently of a preselection.
[0294] As an alternative to the preselection descriptor, preselections may also be defined through the preselection element. The selection of preselections may be based on the included attributes and elements in the preselection element.
[0295] Preselection® id may specify the id of the Preselection. This may be unique within one Period.
[0296] @presel ectionComponents may specify the ids of the included adaptation sets or content components that belong to this preselection as a white space separated list in processing order. The first id may define the main adaptation set.
[0297] ©order - Default: 'undefined' may specify the conformance rules for representations in adaptation sets within the preselection. When set to 'undefined', the preselection may follow the conformance rules for multi-segment tracks. When set to 'time-ordered', the Preselection may follow the conformance rules for time-ordered segment tracks. When set to 'fully-ordered', the preselection may follow the conformance rules for fully-ordered segment tracks. In this case, order in the ©preselectionComponents attribute may specify the component order.
[0298] Accessibility may specify information about an accessibility scheme.
[0299] Role may specify information on role annotation scheme.
[0300] Rating may specify information on a rating scheme.
[0301] Viewpoint may specify information on viewpoint annotation scheme.
[0302] CommonAttributesElements may specify the common attributes and elements (e.g. attributes and elements from base type Representations aseType).
[0303] An example of XML syntax for a Preselection element may be as follows:
<xs:complexType name="PreselectionType"> <xs:annotation>
<xs:documentation xml:lang="en">
Preselection
</x s : documen tation>
</x s : annotation>
<xs:complexContent>
<xs:extension base="RepresentationBaseType">
<xs:sequence>
<xs:element name=" Accessibility' type= " DescriptorT ype ' minOccurs="0" maxOccurs= "unbounded"/>
<xs:element name="Role" type= " DescriptorType " minOccurs="0" maxOccurs= "unbounded"/>
<xs:element name="Rating" type= " DescriptorTy pe " minOccurs="0" maxOccurs= "unbounded"/>
<xs:element name="Viewpoint" type="DescriptorType" minOccurs="0" maxOccurs= "unbounded"/>
</xs:sequence>
<xs:attribute name="id" type="StringNoWhitespaceType" default="l"/>
<xs:attribute name="preselectionComponents" type="StringVectorType" use="required"/>
<xs:attribute name="lang" type="xs:language"/>
<xs:attribute name="order" type="PreselectionOrderType" default="undefined"/>
</xs:extension>
</xs:complexContent>
</x s : complexT ype>
<xs:simpleType name="PreselectionOrderType">
<xs:annotation>
<xs:documentation xml:lang="en">
Preselection Order type
</x s : documen tation>
</x s : annotation>
<xs:restriction base="xs:string">
<xs:enumeration value="undefined"/>
<xs:enumeration value="time-ordered"/>
<xs:enumeration value="fully-ordered"/>
</xs:restriction>
</x s : simpleT ype> [0304] Conformance rules for multi-segment tracks may be specified. Where multiple adaptation sets indicate this type of ordering, each adaptation set and the included representations may follow the regular conformance rules for multi- segment tracks. No additional conformance rules may be defined for the representations in different adaptation sets within preselections.
[0305] Conformance rules for time-ordered segment track may be specified. Where multiple adaptation sets indicate this type of ordering, each adaptation set and the included representations may follow the conformance rules for multi-segment tracks.
[0306] In addition, the concatenation of the following may represent a conforming segment track that also conforms to the media type as specified in the @mimeType attribute for the representation of the main adaptation set: an initialization segment of one representation of the main adaptation set (specified by the first id in the @preselectionComponents attribute or the preselection descriptor; and media segment(s)/subsegment(s) of one representation from each adaptation set referenced in the preselection ordered by non-decreasing first decode time(s). It may be noted that this may not constrain the order of segments with the same first decode time. When adaptation sets within a preselection are time-ordered as defined above, the representations of all adaptation sets referenced by the preselection may be a segment/subsegment.
[0307] Conformance rules for fully-ordered segment track may be specified. Where multiple adaptation sets indicate this type of ordering, each adaptation set and the included Representations may follow the conformance rules for multi-segment tracks.
[0308] In addition, the concatenation of the following may represent a conforming segment track, which may also conform to the media type as specified in the @mimeType attribute for the representation of the main adaptation set: an initialization segment of one representation of the main adaptation set (e.g., specified by the first id in the @preselectionComponents attribute or the preselection descriptor); and media segment(s)/subsegment(s) of one representation from each adaptation set referenced in the preselection, ordered first by non-decreasing decode times and then by position in the list given in @preselectionComponents. When adaptation sets referenced by a preselection are fully ordered as defined above, the representations of all adaptation sets referenced by the preselection may be segment/subsegment aligned.
[0309] Storage of neural network as an item in ISO base media file [0310] WO2022/079545 describes an example method that includes defining a metadata box for a neural network representation (NNR) item data, wherein the NNR item data comprises an NNR bitstream; and defining an association between the NNR item data and an NNR configuration by using a configuration item property, wherein the NNR configuration item property comprises information about stored NNR item data.
[0311] A new item type called ‘NNR item’ may be defined. NNR item may be referred to as NNR item data in some embodiments. The metadata for NNR items may be included in a meta box like the metadata for any other items. NNR items may be stored in an ISOBMFF file or in an external file similarly to any other items as described earlier. For example, NNR items may be stored in the ItemDataBox of the meta box of an ISOBMFF file. Such storage could be at file level, movie box level, or track level. In track level storage, a media track may be dependent on the non-timed neural network to process its samples. The process may be associated to visual enhancement of decoded media data, decoding of the media data in the sample, or alike. In another example, NNR items may be stored in one or more media data boxes (e.g. MediaDataBox) of an ISOBMFF file, while the metadata for items may be included in a meta box, which may be stored at file level, movie box level, or track level.
[0312] ISO/IEC 15938-17 (Compression of Neural Networks for Multimedia Content Description and Analysis) is also known as neural network representation (NNR) or neural network compression (NNC). NNR specifies a compressed representation of the parameters and/or weights of a trained neural network and a decoding process for the compressed representation. NNR complements the description of the network topology in existing neural network exchange formats. NNR is independent of a particular neural network exchange format and is interoperable with common neural network exchange formats.
[0313] NNR establishes a toolbox of compression methods, specifying (where applicable) the resulting elements of the compressed bitstream. All of these tools can be applied to the compression of entire neural networks, and some of them may also be applied to the compression of differential updates of neural networks with respect to a base network. Such differential updates are for example useful when models are redistributed after fine-tuning or transfer learning, or when providing versions of a neural network with different compression ratios. The support for incremental compression of updates of neural networks respective to a base model will be included in the 2nd edition of NNR, which is currently being standardized.
[0314] NNR comprises the syntax format, semantics, associated decoding process requirements, parameter sparsification, parameter transformation methods, parameter quantization, entropy coding method and integration/signalling within existing exchange formats. [0315] FIG. 8 illustrates example structure of a neural network representations (NNR) bitstream. An NNR bitstream may conform to ISO/IEC 15938-17 (Compression of Neural Networks for Multimedia Content Description and Analysis). NNR specifies a high-level bitstream syntax (HLS) for signaling compressed neural network data in a channel as a sequence of NNR Units as illustrated in FIG. 8. As depicted in FIG. 8, according to this structure, an NNR bitstream 802 includes multiple elemental units termed NNR Units (e.g. NNR units 804a, 804b, 804c,... 804n). An NNR Unit (e.g. 804a) represents a basic high-level syntax structure, and includes three syntax elements: NNR Unit Size 806, NNR unit header 808, NNR unit payload 810.
[0316] In some embodiments, NNR bitstream may include one or more aggregate NNR units. An aggregate NNR unit, for example, an aggregate NNR unit 812 may include an NNR unit size 812, an NNR unit header 816, and an NNR unit pay load 818. An NNR unit pay load of aggregate NNR unit may include one or more NNR units. As shown in FIG. 8, the NNR unit payload 818 includes an NNR unit 820a, an NNR unit 820b, and an NNR 820c. The NNR units 820a, 820b, and 820c may include an NNR unit size, an NNR unit header, and an NNR unit payload, each of which may include zero or more syntax elements.
[0317] As mentioned above, a bitstream may be formed by concatenating several NNR Units, aggregate NNR units, or combination thereof. NNR units may include different types of data. The type of data that is included in the payload of an NNR Unit defines the NNR Unit’s type. This type is specified in the NNR unit header. The following table specifies the NNR unit header types and their identifiers.
[0318] NNR unit is data structure for carrying neural network data and related metadata which is compressed or represented using this specification. NNR units carry compressed or uncompressed information about neural network metadata, topology information, complete or partial layer data, filters, kernels, biases, quantization weights, tensors, or the like. An NNR unit may include following data elements:
• NNR unit size: This data element signals the total byte size of the NNR Unit, including the NNR unit size.
• NNR unit header: This data element includes information about the NNR unit type and related metadata.
• NNR unit payload: This data element includes compressed or uncompressed data related to the neural network.
[0319] NNR bitstream is composed of a sequence of NNR Units and/or aggregate NNR units. The first NNR unit in an NNR bitstream shall be an NNR start unit (e.g. NNR unit of type NNR_STR).
[0320] Neural Network topology information can be carried as NNR units of type NNR_TPL. Compressed NN information can be carried as NNR units of type NNR_NDU. Parameter sets can be carried as NNR units of type NNR_MPS and NNR_LPS. An NNR bitstream is formed by serializing these units.
[0321] Image and video codecs may use one or more neural networks at decoder side, either within the decoding loop or as a post-processing step, for both human-targeted and machine targeted compression. [0322] Selected features of the NAL unit file format
[0323] A sample according to ISO/IEC 14496-15 includes one or more length-field-delimited NAL units. The length field may be referred to as NALULength or NALUnitLength. The NAL units in samples do not begin with start codes, but rather the length fields are used for concluding NAL unit boundaries. The scheme of length-field-delimited NAL units may also be referred to as length-prefixed NAL units.
[0324] A VVC subpicture is a rectangular region of one or more slices within a picture. An encoder may treat the subpicture boundaries like picture boundaries and may turn off loop filtering across the subpicture boundaries. Thus, it is possible to encode subpictures so that selected subpictures may be extracted from VVC bitstream(s) or merged to a destination VVC bitstream. Furthermore, such VVC bitstream extraction or merging operations may be performed without modifications of the video coding layer (VCL) NAL units. The subpicture identifiers (IDs) for the subpictures that are present in the bitstream may be indicated in the sequence parameter set(s) or picture parameter set(s).
[0325] ISO/IEC 14496-15 comprises the following types of tracks for carriage of VVC video: VVC track, VVC non- VCL track, VVC subpicture track.
[0326] A VVC track represents a VVC elementary stream by including NAL units in its samples and/or sample entries, and possibly by associating other VVC tracks containing other layers and/or sublayers of the VVC elementary stream through 'vvcb' entity group and the 'vopi' sample group or through the 'opeg' entity group, and possibly by referencing VVC subpicture tracks.
[0327] When a VVC track references VVC subpicture tracks, it is also referred to as a VVC base track. A VVC base track may not contain VCL NAL units and may not be referred to by a VVC track through a 'vvcN' track reference.
[0328] versatile video coding (VVC) non-video coding layer (non-VCL) track (VVC non-VCL) track is a track that includes only non-VCL NAL units and is referred to by a VVC track through a 'vvcN' track reference.
[0329] A VVC subpicture track includes either of the following: a sequence of one or more VVC subpictures forming a rectangular region, or a sequence of one or more complete slices forming a rectangular region. A sample of a VVC subpicture track contains either of the following: one or more complete subpictures that form a rectangular region, or one or more complete slices that form a rectangular region. [0330] VVC subpicture tracks enable the storage of VVC subpictures as separate tracks so that any combination of subpictures can be streamed or decoded. VVC subpicture tracks enable representing rectangular regions of the same video content at different bitrates or resolutions. Consequently, bitrate or resolution emphasis on regions can be adapted dynamically by selecting the VVC subpicture tracks that are streamed or decoded.
[0331] When a VVC subpicture track contains a conforming VVC bitstream, which may be consumed without other VVC subpicture tracks, a regular VVC sample entry is used ('vvcl' or 'vvil') for the VVC subpicture track. When a VVC subpicture track does not contain a conforming VVC bitstream, a 'vvsl' sample entry is used for the VVC subpicture track.
[0332] A file reader may reconstruct a VVC bitstream, for example, when a VVC track has an associated VVC non-VCL track. In implicit reconstruction of a VVC bitstream, an access unit is reconstructed from the time-aligned samples of the VVC track and the associated VVC non-VCL track having a current decoding time so that the NAL unit order conforms to VVC.
[0333] Storage of NNR data as NAL units in a separate track
[0334] WO2022/079545 also describes storage of NNR coded data as NAL units in a track that is linked to a media track. Some examples are provided below:
[0335] In an example, NN tracks could have time-aligned samples with the samples of associated media tracks.
- Global NN weight updates could be stored e.g. in a new box in the track-level, a sample entry, sample group entries, or samples. Storing them in sample entry or sample group entry may be beneficial over the other alternatives e.g. for unicast streaming services.
- NN weight updates may be aligned with the Group of Picture structures of media data and media track. In such a case the NN data could be stored in the samples of NN track.
[0336] In an example, NNR coded data may be stored in conjunction with media data as follows:
- There could be NAL units which are NNR compressed and defined in audio/video codec scope.
Such NAL units could be collected in NN tracks as samples of this track then linked to the video track The same NN track may be linked or referenced by multiple audio/video tracks. This could be especially useful when there are multiple representations of the same media track which can utilize the same NN in its samples.
[0337] In another example, such media samples may include further NN updates in their bitstreams as specially marked data structures (e.g., NN specific NAL units).
[0338] In another example, there could be NAL units which are NNR compressed and defined in an audio/video codec context. Such NAL units may be aggregated in NNR tracks as samples of NNR track and then linked to the media tracs that they apply. Time alignment of such media tracks may be done via the sample composition timing mechanisms.
[0339] In an example, WO2022/079545 describes a method including defining an NNR track sample including one or more NNR units. The NNR track sample is linked to an NNR unit parameter via one or more of a sample entry, a sample group, or a non-timed NNR item data. In another example, the method further includes defining a network abstraction layer (NAL) unit, which is aggregated in NNR tracks as a sample of NNR track and linked to the associated media track. In an example, the NNR sample track is may be independently decodable or dependent of previous sample for decoding. The NNR sample track may be used by other media tracks for decoding media samples of the other media tracks. A track referencing mechanism may be used to associate the NNR sample track with the other media tracks for decoding the media samples of the other media tracks.
[0340] The method may further include defining a network abstraction layer (NAL) unit, wherein the NAL unit is aggregated in NNR tracks as a sample of NNR track and linked to the associated media track.
[0341] Usage of SEI prefix indication SEI message with post-filter SEIs for media presentation characteristics signaling
[0342] An entity makes one or more SEI prefix indications available along a video bitstream, e.g. in a media description, such as SDP or DASH MPD. An SEI prefix indication may comprise a SEI prefix indication SEI message or may comprise an initial part or an entire syntax structure of one or more SEI messages, such as a post-filter related SEI messages. An SEI prefix indication may be, but is not limited to, one or more of the following:
• One or more MIME media parameters. The MIME media parameter(s) may be encapsulated in an SDP parameter or in an attribute of a streaming manifest (e.g. DASH MPD) or alike. Separate or same MIME media parameter(s) may be used for declarative bitstream properties, encoding capabilities, and/or preferences or requirements for bitstream to be decoded.
• One or more attributes, parameters, or alike of the media description. An attribute may for example be an attribute in DASH MPD.
• One or more syntax elements or syntax structures in a file encapsulating a bitstream.
[0343] Since NNPFC SEI messages can be big in size and since it is optional for players and decoders to support the NNPFC SEI message, there would need to be mechanisms to indicate the presence and/or type of NNPFC SEI messages for a track (in ISOBMFF) or for a representation (in DASH). Consequently, the player or decoder may determine whether to process a track or representation including NNPFC SEI messages.
[0344] SEI NAL units including NNPFC and/or NNPFA SEI messages can be stored in an ISO base media file similarly to any other NAL units. However, such a storage has at least the following drawbacks as example:
NNPFC SEI messages have a scope of a coded video sequence, and therefore need to be re-included in the bitstream for each coded video sequence. The current mechanisms lack the capability of avoiding transmitting NNPFC SEI messages with the same content repeatedly.
Loading an NN model to the inference engine (e.g., a graphics processing unit) may take time. The current mechanisms for carrying NNPFC and NNPFA SEI messages do not facilitate a lookahead mechanism to prepare for loading the NN model.
[0345] In an embodiment, a file writing method is provided. The file writing method includes: storing coded video data in a video track of a file; storing neural-network post-filter (NNPF) supplemental enhancement information (SEI) in the video track or in a second track associated with the video track; indicating, in the file, the presence of the NNPF SEI in the video track or the second track with at least one of a scheme type of a restricted video sample entry type and an essential sample group for the NNPF SEI. NNPF SEI may be defined as a collective term for NNPFC SEI message(s) and/or NNPFA SEI message(s) and/or SEI NAL unit(s) including NNPFC SEI message(s) and/or NNPFA SEI message(s).
[0346] In another embodiment, a file parsing method is provided. The file parsing method includes: reading, from a file, the presence of the NNPF SEI in a video track or a second track from at least one of a scheme type of a restricted video sample entry type and an essential sample group for the NNPF SEI; in response to supporting NNPF SEI, processing the NNPF SEI in the video track or the second track; decoding coded video data from the video track; filtering the decoded video data with a post-filter defined by the NNPF SEI.
[0347] Both of the following approaches are possible for carriage of NNPFC and NNPFA SEI messages:
NNPFC and NNPFA SEI messages are included in a track that also includes coded video data, o There might be multiple tracks with the same coded video data. One track might not include NNPFC and NNPFA SEI messages and is intended for players and decoders that do not support NNPFC and NNPFA SEI messages. Other tracks may include or reference different post-filters.
NNPFC and NNPFA SEI messages are included in a track that is separate from the coded video data. For example, NNPFC and NNPFA SEI messages are included in a VVC non-VCL track. Players and decoders need to access the track with the coded video data and additionally players and decoders may access the track with the NNPFC and NNPFA SEI messages.
[0348] Restricted media scheme type to indicate post-decoder requirement to process NNPFC SEI messages
[0349] In an embodiment, a restricted media scheme type is defined to indicate that the track requires handling of NNPFC and NNPFA SEI messages for post-processing. In addition, a restricted media scheme type may entail limitations on NNPFC and/or NNPFA SEI messages in the track, such as limitations of the allowed syntax element values of NNPFC SEI messages. Examples of such restricted media scheme type definitions are described below.
[0350] A specific scheme type, such as 'nnfa', indicates that the track complies with the 'aSEI' scheme and requires the processing of NNPFC and NNPFA SEI messages for post-processing.
[0351] A specific scheme type, such as 'nnfO', indicates that the track complies with the 'aSEI' scheme and requires the processing of NNPFC SEI messages with nnpfc_mode_idc equal to 0 for indicating base post-processing filters and nnpfc_mode_idc equal to 1 for filter updates and NNPFA SEI messages for post-processing. The base filter that is provided to the decoding or post-processing may be identified, for example, as described in the embodiment ‘Item property for associating an nnpfc_id value to a neural network item in ISOBMFF’ . [0352] A specific scheme type, such as 'nnfl', indicates that the track complies with the 'aSEI' scheme and requires the processing of NNPFC SEI messages with nnpfc_mode_idc equal to 1 (indicating that an ISO/IEC 15398-17 bitstream is included in the SEI message) and NNPFA SEI messages for post-processing.
[0353] A specific scheme type, such as 'nnf2', indicates that the track complies with the 'aSEI' scheme and requires the processing of NNPFC SEI messages with nnpfc_mode_idc equal to 2 for indicating base post-processing filters and nnpfc_mode_idc equal to 1 for filter updates and NNPFA SEI messages for post-processing.
[0354] Both 'aSEI' and any scheme type described above can be indicated for the same track. One scheme type can be included in the SchemeTypeBox and another can be indicated in a CompatibleSchemeTypeBox.
[0355] As specified in ISOBMFF, the scheme types of a track can be included in the codecs MIME parameter as a list of four-character codes separated by a plus ('+') character. For example, the codecs parameter for a VVC track that requires processing of NNPFC and NNPFA SEI messages could be codecs=’resv.aSEI+nnfa.vvcl’. In another example, the codecs parameter for a VVC non-VCL track that requires processing of NNPFC and NNPFA SEI messages could be codecs=’resv.aSEI+nnfa.vvcN’ .
[0356] The codecs MIME parameter may be used in a media presentation description to indicate the properties of a track or Representation.
[0357] Indicating parameters of the neural network
[0358] It may be beneficial for a player or reader or decoder to get to know parameters of the neural network, such as complexity parameters, to determine when is possible to use it. Embodiments for indicating parameters of the neural network are described below.
[0359] In an embodiment, a neural-network post-filter information box (e.g., with 4CC equal to 'nnfi') is defined to carry syntax elements of an NNPFC SEI message, but without NNR bitstream, or any information similar to what is carried in an NNPFC SEI message. The container for the neural- network post-filter information box may be the SchemelnformationBox. The syntax may be specified as follows: aligned(8) class NNPFInformationBox extends FullBoxfnnfi',0,0) { unsigned int(8) nnpfc_sei_data_byte[];
}
[0360] The byte array nnpfc_sei_data_byte[ ] includes a bit string that starts an NNPFC SEI message.
[0361] In an embodiment, a SEI prefix indication box (e.g. with 4CC equal to 'seip') is defined to include a SEI prefix indication SEI message. The container for the SEI prefix indication box may be the SchemelnformationBox. The syntax may be specified as follows: aligned(8) class SEIPrefixIndicationBox extends FullBox('seip',0,0) { unsigned int(8) sei_prefix_sei_data_byte[];
}
[0362] The byte array sei_prefix_sei_data_byte[ ] includes a SEI prefix indication SEI message.
[0363] It may be advantageous for a player or reader or decoder to get to know the order of SEI- defined post-processing steps to determine when it is possible to perform the intended post-processing. In an embodiment, a SEI processing order box (e.g., with 4CC equal to 'seio') is defined to include a SEI processing order SEI message. The container for the SEI prefix indication box may be the SchemelnformationBox. The syntax may be specified as follows: aligned(8) class SEIProcessingOrderBox extends FullBox('seio',0,0) { unsigned int(8) sei_processing_order_sei_data_byte[];
}
[0364] The byte array sei_processing_order_data_byte[ ] includes a SEI processing order SEI message.
[0365] Sample group to include NNPFC SEI messages
[0366] In an embodiment, an NNPFC SEI sample group is specified as described in the following paragraphs. [0367] The NNPFC SEI sample group may be indicated to be an essential sample group. Consequently, a player is required to process the NNPFC SEI sample group. For example, a player processes the NNPFC SEI sample group for bitstream reconstruction. When an NNPFC SEI sample group is not indicated to be an essential sample group, the player may ignore processing of the NNPFC SEI sample group.
[0368] When an essential sample group is used, the NNPFC SEI sample group type (e.g., 'nfcs') may be included in the essential MIME parameter, which may be included in a media description, such as in the @mimeType parameter of DASH MPD.
[0369] When more than one NNPFC SEI message with the same nnpfc_id value is present in a track, the SampleToGroupBox(es) for this sample group type may include grouping_type_parameter equal to nnpfc_id.
[0370] The syntax of the sample group description entry may be specified as follows: class NnpfcSeiEntryO extends VisualSampleGroupEntry fnfcs'){ unsigned int(8) nnpfc_sei_data_byte[];
}
[0371] The byte array nnpfc_sei_data_byte[ ] includes an NNPFC SEI message.
[0372] Example embodiments
[0373] In an embodiment following two sample groups are proposed:
NNPFC sample group, where a sample group description entry may be referred to as NnpfcSeiEntry and includes an NNPFC SEI message.
NNPFA sample group, where a sample group description entry includes an NNPFA SEI message.
[0374] Some of the features of the proposed sample groups is as follows:
The NNPFC SEI message has a persistence scope of a coded layer video sequence (CLVS). However, the same post-filter may be used for many CLVSs. Since NNPFC SEI messages may include NNR bitstreams, they may be of substantial size. The NNPFC sample group avoids repetitive presence of NNPFC SEI messages in tracks and thus reduces file size. File writers may make NNPFC and NNPFA sample groups essential when the postprocessing filtering is mandatory. When NNPFC and NNPFA essential sample groups are present in a VVC non-VCL track, the processing of the entire track may be omitted when the reader is not capable of post-filtering.
When NNPFC SEI messages are represented by a non-essential NNPFC sample group and a reader does not process the sample group, the bitstream does not include NNPFC SEI messages and consequently must not include NNPFA SEI messages either. Thus, an NNPFA sample group is proposed as a companion to the NNPFC sample group, so that a reader either processes both or skips both. In addition, the NNPFA sample group avoids repetitive presence of NNPFA SEI messages in tracks.
[0375] A base post-processing filter typically has a longer persistence than any of the updates applied on top of the base post-processing filter, the NNPFC sample group uses a flag in grouping_type_parameter to indicate whether SampleToGroupBox includes a mapping of the base postprocessing filters or a mapping of filter updates. Additionally, grouping_type_parameter includes filter_id (i.e. , nnpfc_id), since the persistence of updates for different filters are not necessarily aligned and hence included in different instances of SampleToGroupBox.
[0376] FIGs. 9a, 9b, and 9c illustrate an example of the proposed sample groups, in accordance with an embodiment. In this example, two base post-processing filters (IDs 0 and 1) are defined. Two updates on top of each base post-processing filter are provided, labeled as updates A and B on top of base post-processing filter 0 and updates C and D on top of base post-processing filter 1. Both base post-processing filters persist for the entire duration of the track (Z samples). The updates A, B, C, and D have persistence of a, b, c, and d samples, respectively. Two NNPFA SEI messages are defined in the sample group description of the NNPFA sample group, one activating filter ID 0 and another activating filter ID 1. The first three entries in the SampleToGroupBox of the NNPFA sample group activates filter IDs 1, 0, and 1 for x, y, and z samples, respectively, as further detailed below:
Since c > x, the first entry activates the post-processing filter that is obtained by applying the update C on top of the base post-processing filter 1.
The second entry activates the post-processing filter that is obtained by applying the update A on top of the base post-processing filter 0 for an initial part of the period of y samples until the filter update B takes place. After that, the second entry activates the post-processing filter obtained by applying the update B on top of the base post-processing filter 0 for the latter part of the period of y samples.
The third entry activates the post-processing filter that is obtained by applying the update D on top of the base post-processing filter 1. [0377] An NNPFC SEI message includes the nnpfc_id syntax element, which is an identifying number that may be used to identify the post-processing filter that the NNPFC SEI message concerns.
[0378] An NNPFC SEI message identifies an applicable post-processing filter associated with the nnpfc_id value. The use of applicable post-processing filters with different values of nnpfc_id for specific pictures is indicated with neural-network post-filter activation (NNPFA) SEI messages.
[0379] An NNPFC SEI message either specifies a base post-processing filter or includes a neural network update. A base post-processing filter is identified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within a CLVS. When there is no subsequent NNPFC SEI message that has the same nnpfc_id value as the base post-processing filter, the applicable postprocessing filter is the same as the base post-processing filter. Otherwise, the applicable post-processing filter is obtained by applying the update provided as an ISO/IEC 15938-17 bitstream in a subsequent NNPFC SEI message on top of the base post-processing filter.
[0380] Instances of the SampleToGroupBox for the NNPFC sample group include grouping_type_parameter. The grouping_type_parameter field is specified for the NNPFC sample group as follows:
{ unsigned int(l) filter_update_flag; unsigned int(31) filter_id;
}
[0381] filter_update_flag equal to 1 indicates that all the sample group description entries referenced by this SampleToGroupBox include an NNPFC SEI message that has nnpfc_mode_idc equal to 1 and provides an update on top of abase post-processing filter. filter_update_flag equal to 0 indicates that the sample group description entries referenced by this SampleToGroupBox includes an NNPFC SEI message that specifies a base post-processing filter.
[0382] filter_id indicates that the sample group description entries referenced by this SampleToGroupBox include an NNPFC SEI message that has nnpfc_id equal to filter_id.
[0383] Accordingly, the grouping_type_parameter definition, the post-processing filters for different nnpfc_id values are specified in different instances of the SampleToGroupBox. Further, one SampleToGroupBox specifies the base post-processing filter(s) for a particular nnpfc_id value, while another SampleToGroupBox, when any, specifies the filter updates for the same nnpfc_id value. It is therefore possible to indicate that the base post-processing filter persists over a longer period than any of the filter updates.
[0384] When a sample is not mapped to NnpfcSeiEntry in a SampleToGroupBox having filter_update_flag equal to 0 and a particular filter_id, the sample is not be mapped to an NnpfcSeiEntry in a SampleToGroupBox having filter_update_flag equal to 1 and the same filter_id.
[0385] Requirements for a reader to process an NNPFC sample group are specified as follows:
— When the NNPFC sample group is an essential sample group that is present in a VVC non-VCL NAL unit and a reader does not recognize the NNPFC sample group, the reader ignores and skips the VVC non-VCL track in the bitstream reconstruction.
— Otherwise, when the NNPFC sample group is an essential sample group, readers perform the following implicit insertion of prefix SEI NAL units, while otherwise the following implicit insertion of prefix SEI NAL units is optional:
• When a sample is mapped to at least one NnpfcSeiEntry and the sample is either a sync sample or the first sample of a sequence of samples associated with the same sample entry, the sample implicitly includes a prefix SEI NAL unit for each layer included in the track and each filter_id value mapped to the sample, and the prefix SEI NAL unit includes the NNPFC SEI message from the NnpfcSeiEntry with filter_update_flag equal to 0, followed by the NNPFC SEI message from the NnpfcSeiEntry with filter_update_flag equal to l,if any.
[0386] Syntax aligned(8) class NnpfcSeiEntryO extends VisualSampleGroupEntry('nfcs') { unsigned int(8) nnpfc_sei_data_byte[];
}
[0387] Semantics
[0388] nnpfc_sei_data_byte[] is a byte array that includes exactly one complete NNPFC SEI message. [0389] In an alternative embodiment, an NNPFC SEI sample group description entry includes a prefix SEI NAL unit that includes one or more NNPFC SEI messages with the same nnpfc_id value.
[0390] A prefix SEI NAL unit with NNPFC SEI message(s) from a sample group description entry is implicitly included in the first sample of each run of samples mapped to this sample group description entry and in each sync sample within each run of samples mapped to this sample group description entry.
[0391] In an embodiment, it is indicated whether a sample group includes NNPFC SEI messages defining the base post-processing filter or NNPFC SEI messages defining updates of the base postprocessing filter. In different embodiments, this functionality is enabled by one of the following: grouping_type_parameter includes the information whether the SampleToGroupBox includes NNPFC SEI messages defining the base post-processing filter or NNPFC SEI message defining updates of the base post-processing filter. For example, the most significant bit of grouping_type_parameter may be used for this purpose, while the remaining bits define the nnpfc_id value.
Different grouping type four-character codes may be used for the base post-processing filter and updates, such as 'ncbf and 'ncfu', respectively.
[0392] In an embodiment, a base post-processing filter is provided through means other than a sample group description entry, including, but not limited to, any of the following:
An item including a neural network; and
A box directly or indirectly included in a sample entry.
[0393] Sample group to include NNPFA SEI messages
[0394] In an embodiment, an NNPFA SEI sample group is specified as described in the following paragraphs.
[0395] The NNPFA SEI sample group may be indicated to be an essential sample group. Consequently, a player is required to process the NNPFA SEI sample group.
[0396] It may be specified that grouping_type_parameter is absent for the SampleToGroupBox of a NNPFA SEI sample group. Alternatively, it may be specified that grouping_type_parameter indicates the layer or layers similarly to what is done for other sample groups, such as 'sap '. [0397] The syntax of the sample group description entry may be referred to as NnpfaSeiEntry and may be specified as follows: class NnpfaSeiEntryO extends VisualSampleGroupEntry fnfas'){ unsigned int(8) nnpfa_sei_data_byte[];
}
[0398] The byte array nnpfa_sei_data_byte[ ] includes an NNPFA SEI message.
[0399] The NNPFA SEI message enables to indicate efficiently the same NNPFA SEI message being used for multiple samples, since a single NNPFA SEI message may be mapped to a run of samples in a SampleToGroupBox.
[0400] An NNPFA SEI message includes the nnpfa_id syntax element, which is an identifying number that may be used to identify the post-processing filter that the NNPFA SEI message concerns.
[0401] An NNPFA SEI message indicates that the applicable post-processing filter with nnpfc_id equal to nnpfa_id may be used to filter the picture including the NNPFA SEI message.
[0402] Requirements for a reader to process an NNPFA sample group are specified as follows:
— When the NNPFA sample group is an essential sample group that is present in a VVC non-VCL NAL unit and a reader does not recognize the NNPFA sample group, the reader ignores and skips the VVC non-VCL track in the bitstream reconstruction.
Otherwise, when the NNPFA sample group is an essential sample group, readers perform the following implicit insertion of prefix SEI NAL units, while otherwise the following implicit insertion of prefix SEI NAL units is optional:
When a sample is mapped to at least one NnpfaSeiEntry, the sample implicitly includes a prefix SEI NAL unit for each layer included in the track, and the prefix SEI NAL unit includes the NNPFA SEI message from the NnpfaSeiEntry.
[0403] When a reader processes an NNPFA sample group, it also processes the NNPFC sample groups of the same track.
[0404] When an NNPFC sample group is an essential sample group and an NNPFA sample group is present in the same track, the NNPFA sample group is an essential sample group and the 'esgh' sample group lists 'ncfs' and 'ncfa' in subsequent entries of the sample_group_description_type array. [0405] Syntax aligned(8) class NnpfaSeiEntryO extends VisualSampleGroupEntry('nfas')
{ unsigned int(8) nnpfa_sei_data_byte[];
}
[0406] Semantics nnpfa_sei_data_byte[] is a byte array that includes exactly one complete NNPFA SEI message.
[0407] In an alternative embodiment, an NNPFA SEI sample group description entry includes a prefix SEI NAL unit that includes one NNPFA SEI message.
[0408] A prefix SEI NAL unit with NNPFA SEI message from a sample group description entry is implicitly included in each sample mapped to this sample group description entry.
[0409] Sample group to include any SEI messages, with mapping to NNPFC SEI messages
[0410] In an embodiment, an SEI message sample group is specified as described in the next paragraphs.
[0411] The SEI message sample group may be indicated to be an essential sample group. Consequently, a player is required to process the SEI sample group.
[0412] grouping_type_parameter is defined to have 16 bits of SEI type (e.g., in the most significant bits) and 16 bits (e.g. in the least significant bits) that may be specified for specific SEI type. In other words, grouping_type_parameter may have the following syntax breakdown:
{ unsigned int(16) sei_payload_type; unsigned int(16) payload_type_specific_parameter;
}
[0413] In the case of NNPFC SEI, payload_type_specific_parameter may be required to be unique for each nnpfc_id value being used. [0414] In the case of NNPFC SEI, values of payload_type_specific_parameter may be specified through a SampleGroupDescriptionBox with a specific grouping_type, such as 'nfim', whose sample group description provides a mapping between values of payload_type_specific_parameter and nnpfc_id. Such a sample group description entry may, for example, have the following syntax: class NnpfcIdMappingEntryO extends VisualSampleGroupEntry ('nfim'){ unsigned int(16) payload_type_specific_parameter; unsigned int(32) nnpfc_id;
}
[0415] The SampleGroupDescriptionBox including NnpfcIdMappingEntry structures may be used to find which SEI message sample group description entry includes an NNPFC SEI message with a nnpfc_id value equal to a nnpfa_id value used in a NNPFA SEI message.
[0416] SampleToGroupBox for 'nfim' is not present.
[0417] When SampleGroupDescriptionBox including NnpfcIdMappingEntry structures is not present, payload_type_specific_parameter may be required to be equal to nnpfc_id.
[0418] The syntax of the SEI message sample group description entry may be specified as follows: class SeiEntryO extends VisualSampleGroupEntry ('seim'){ unsigned int(8) persistence_scope; unsigned int(8) sei_data_byte[];
}
[0419] persistence_scope equal to 0 indicates that the SEI message applies to the current sample only. persistence_scope equal to 1 indicates that the SEI message applies up to the next sync sample (exclusive). Other persistence_scope values are reserved.
[0420] The byte array sei_data_byte[ ] includes an SEI message.
[0421] At least following variations of this embodiment are possible:
One sample group type is defined for SEI messages carried in prefix SEI NAL units, and another sample group type is defined for SEI messages carried in suffix SEI NAL units. Rather than just an SEI message, the sample group description entry includes an SEI NAL unit (which may be constrained to include one SEI message, or may be constrained to include one or more SEI messages of the same payload type, or may include one or more SEI messages of any payload type).
In an example, 'psei' sample group description entry carries one or more prefix SEI NAL units of the same payload type, and 'ssei' sample group description entry carries one or more suffix SEI NAL units of the same payload type. grouping_type_parameter may be defined as above.
[0422] When persistence_scope is equal to 0, a prefix SEI NAL unit with SEI message(s) from a sample group description entry is implicitly included in each sample mapped to this sample group description entry.
[0423] When persistence_scope is equal to 1, a prefix SEI NAL unit with SEI message(s) from a sample group description entry is implicitly included in the first sample of each run of samples mapped to this sample group description entry and in each sync sample within each run of samples mapped to this sample group description entry.
[0424] In general, any sample group described above may be indicated to be an essential sample group. Consequently, a player is required to process the sample group, e.g., for bitstream reconstruction. Likewise, when a post-filter is regarded as optional by file writer, related sample groups may be indicated to be non-essential. Alternatives for indicating mandatory processing of a sample group in (VVC) non-VCL track include:
A new track reference from a VVC track to the VVC non-VCL track, which indicates mandatory processing of selected above-described sample groups.
A new sample entry type for a VVC non-VCL track, which indicates mandatory processing of selected above-described sample groups.
A new restricted media scheme type (e.g., as described above), which indicates mandatory processing of selected above-described sample groups.
[0425] It is remarked that in the indications above, the indications may be pre-defined to indicate either of the following two cases or may be capable of indicating which of the following two cases applies:
Selected above-described sample groups must be processed to play a bitstream. When they are not processed, the video bitstream must not be decoded or played. Selected above-described sample groups must be processed to process the non-VCL track including them. When they are not processed, the video bitstream reconstructed without the non-VCL track can be decoded or played.
[0426] Sample group to include NNPFC and NNPFA SEI messages
[0427] In an embodiment, the NNPFC SEI and the NNPFA SEI NAL units may be signaled/specified together in the same sample group as described in the following paragraphs.
[0428] In an example embodiment the sample group carrying both NNPFC SEI and the NNPFA SEI NAL units together may be called the NNPF SEI sample group.
[0429] The NNPF SEI sample group may be indicated to be an essential sample group. Consequently, a player is required to process the NNPF SEI sample group. For example, a player processes the NNPF SEI sample group for bitstream reconstruction. When an NNPF SEI sample group is not indicated to be an essential sample group, the player may ignore processing of the NNPF SEI sample group.
[0430] When an essential sample group is used, the NNPF SEI sample group type (e.g., 'nfcs') may be included in the essential MIME parameter, which may be included in a media description, such as in the @mimeType parameter of DASH MPD.
[0431] When more than one NNPFC SEI message with the same nnpfc_id value is present in a track, the SampleToGroupBox(es) for this sample group type may include grouping_type_parameter equal to nnpfc_id.
[0432] In an altetnate embodiment, it may be specified that grouping_type_parameter is absent for the SampleToGroupBox of a NNPF SEI sample group. In another alternate embodiment, it may be specified that grouping_type_parameter indicates the layer or layers similarly to what is done for other sample groups, such as 'sap '.
[0433] The syntax of the sample group description entry may be specified as follows: class NnpfSeiEntryO extends VisualSampleGroupEntry ('npfs'){ unsigned int(16) num_nalus; for (i=0; i< num_nalus; i++) { unsigned int (8) nal_unit_type; unsigned int(16) nal_unit_length; bit(8*nal_unit_length) nal_unit;
}
}
[0434] nal_unit_type indicates the type of the NAL units in the sample group; it takes a value as defined in ISO/IEC 23090-3; it is restricted to take one of the values indicating a prefix SEI NAL unit (e.g., NNPFC SEI and NNPFA SEI).
[0435] num_nalus indicates the number of NAL units included in the sample group.
[0436] nal_unit_length indicates the length in bytes of the NAL unit.
[0437] nal_unit includes a declarative SEI NAL unit (e.g., NNPFC SEI and NNPFA SEI), as specified in ISO/IEC 23090-3.
[0438] Example embodiments
[0439] In an embodiment following NNPF SEI sample group is proposed:
NNPF SEI sample group, where a sample group description entry may be referred to as NnpfSeiEntry and includes both NNPFC SEI message and the corresponding NNPFA SEI message.
[0440] The NNPFA SEI message enables to indicate efficiently the same NNPFA SEI message being used for multiple samples, since a single NNPFA SEI message may be mapped to a run of samples in a SampleToGroupBox.
[0441] An NNPFA SEI message includes the nnpfa_id syntax element, which is an identifying number that may be used to identify the post-processing filter that the NNPFA SEI message concerns.
[0442] An NNPFA SEI message indicates that the applicable post-processing filter with nnpfc_id equal to nnpfa_id may be used to filter the picture including the NNPFA SEI message.
[0443] Instances of the SampleToGroupBox for the NNPF sample group include grouping_type_parameter. The grouping_type_parameter field is specified for the NNPF sample group as follows: { unsigned int(l) filter_update_flag; unsigned int(31) filter_id;
}
[0444] filter_update_flag equal to 1 indicates that all the sample group description entries referenced by this SampleToGroupBox include an NNPFC SEI message that has nnpfc_mode_idc equal to 1 and provides an update on top of abase post-processing filter. filter_update_flag equal to 0 indicates that the sample group description entries referenced by this SampleToGroupBox includes an NNPFC SEI message that specifies a base post-processing filter.
[0445] filter_id indicates that the sample group description entries referenced by this SampleToGroupBox include an NNPFC SEI message that has nnpfc_id equal to filter_id.
[0446] Accordingly, the grouping_type_parameter definition, the post-processing filters for different nnpfc_id values are specified in different instances of the SampleToGroupBox. Further, one SampleToGroupBox specifies the base post-processing filter(s) for a particular nnpfc_id value, while another SampleToGroupBox, when any, specifies the filter updates for the same nnpfc_id value. It is therefore possible to indicate that the base post-processing filter persists over a longer period than any of the filter updates.
[0447] When a sample is not mapped to NnpfSeiEntry in a SampleToGroupBox having filter_update_flag equal to 0 and a particular filter_id, the sample is not be mapped to an NnpfSeiEntry in a SampleToGroupBox having filter_update_flag equal to 1 and the same filter_id.
[0448] Requirements for a reader to process an NNPFC sample group are specified as follows:
— When the NNPF sample group is an essential sample group that is present in a VVC non-VCL NAL unit and a reader does not recognize the NNPF sample group, the reader ignores and skips the VVC non-VCL track in the bitstream reconstruction.
— Otherwise, when the NNPF sample group is an essential sample group, readers perform the following implicit insertion of prefix SEI NAL units, while otherwise the following implicit insertion of prefix SEI NAL units is optional:
• When a sample is mapped to at least one NnpfSeiEntry and the sample is either a sync sample or the first sample of a sequence of samples associated with the same sample entry, the sample implicitly includes a prefix SEI NAL unit for each layer included in the track and each filter_id value mapped to the sample, and the prefix SEI NAL unit includes the NNPFC SEI message from the NnpfSeiEntry with filter_update_flag equal to 0, followed by the NNPFC SEI message from the NnpfSeiEntry with filter_update_flag equal to l,if any.
[0449] Bitstream reconstruction from sample groups
[0450] SEI NAL unit(s) from a sample group description entry that persists until the next sync sample are included in an access unit reconstructed from the first sample subject to bitstream reconstruction or a sync sample. SEI NAL unit(s) from a sample group description entry that persists for the current sample are included in the respective access unit.
[0451] When an NNPFC SEI message includes a base post-processing filter and another NNPFC SEI message includes a filter update for the same nnpfc_id value and both these SEI messages are to be included in the same access unit, they are included in the same SEI NAL unit.
[0452] SEI NAL unit(s) including SEI message(s) may be implicitly included from the sample group description entry in the sample as described above, or a bitstream reconstruction process may explicitly include SEI NAL unit(s) from sample group description entries in the bitstream. In an example embodiment where 'psei' sample group may be present in a VVC non-VCL track, the following steps 1 and 2 may be added to the bitstream reconstruction process (subclause 11.6.3 of ISO/IEC 14496- 15 6th Edition).
[0453] When there is an associated VVC non-VCL track and the picture unit is the first picture unit in this access unit that is reconstructed from the sample, the following applies:
1. For each 'psei' sample group present in the associated VVC non-VCL track, when the time-aligned sample of the associated VVC non-VCL track is mapped to a sample group description entry that has persistence_scope equal to 1 and the current sample is a sync sample or the first sample of a sequence of samples associated with the same sample entry, the SEI NAL unit(s) included in the sample group description entry are included in the picture unit.
2. For each 'psei' sample group present in the associated VVC non-VCL track, when the time-aligned sample of the associated VVC non-VCL track is mapped to a sample group description entry that has persistence_scope equal to 0, the SEI NAL unit(s) included in the sample group description entry are included in the picture unit.
3. The following NAL units are included in the picture unit: o When there is at least one NAL unit in the time-aligned sample of the associated VVC non-VCL track with nal_unit_type equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, or RSV_NVCL_27 (a NAL unit with such a NAL unit type cannot precede the first VCL NAL unit in a picture unit), the NAL units (excluding the AUD NAL unit, if any) in the time-aligned sample of the associated VVC non-VCL track up to and excluding the first of these NAL units. o Otherwise all NAL units in the time-aligned sample of the associated VVC non-VCL track.
[0454] Item property for associating an nnpfc id value to a neural network item in ISOBMFF
[0455] A neural network may be stored as an item. The neural network data may, for example, comply with NNR. The item data may reside in a standalone file, which may be identified with a URL. The URL may be provided in a DataEntryUrlBox as an entry in the DataReferenceBox, and the data reference entry may then be associated with an item in an ItemLocationBox. Alternatively, the item data may reside in the ISO base media file, for example in ItemDataBox.
[0456] When the neural network data resides in a standalone file which is identified by a URL, a client needs to fetch the resource pointed to be the URL only once, as opposed to fetching it again and again, e.g., in the case of storing the neural network data in the video bitstream.
[0457] In an embodiment, a specific handler, such as 'nrnw', that is associated with an item including a neural network (for example, through HandlerBox or HandlerProperty) may indicate that the item comprises a neural network. The item type may indicate the format of the item. For example, the item type 'nnrl' or 'nncl' may indicate that the item data conforms to ISO/IEC 15938-7.
[0458] In an embodiment, a new item property is defined to provide an identifier value, such as an nnpfc_id value, which associates the neural network stored in the item data to an identifier that may be used to associate NN characteristics provided in a video bitstream (e.g. in a NNPFC SEI message) and/or activate a NN for one or more pictures in a video bitstream (e.g. through a NNPFA SEI message). This item property may be present for a neural network item that is present in a track that includes NNPFC and/or NNPFA SEI messages. The item property may have the following syntax: aligned(8) class NNPostFilterProperty(property_type, version, flags) extends ItemFullProperty('nnfp', 0, 0)
{ unsigned int(32) nnpf_id;
}
[0459] nnpf_id indicates that the track includes an NNPFC SEI message with nnpfc_mode_idc equal to 0 (the base post-processing filter associated with the nnpfc_id value is determined by external means) and nnpfc_id is equal to nnpf_id. The neural network included in the item data serves as the base post-processing filter determined by external means.
[0460] In an embodiment, NNPostFilterProperty may comprise syntax elements of NNPFC SEI message or any information similar to what is included theNNPFC SEI message.
[0461] Usage of a preselection with VVC tracks and VVC non-VCL tracks
[0462] In an embodiment, the VVC tracks and VVC non-VCL tracks including NNPFC and NNPFA SEI messages may be part of a preselection and are intended to be consumed jointly. In an embodiment, the VVC track forms a representation in the main adaptation set, and VVC non-VCL track forms a Representation in the partial adaptation set.
[0463] In an embodiment, the presence of tracks including NNPFC and NNPFA SEI messages in preselection may be indicated in PreselectionGroupBox and/or PreselectionProcessingBox and/or PreselectionTrackGroupEntryBox.
[0464] In an embodiment, the PreselectionGroupBox may comprise neural-network post-filter information NNPFInformationBox, as described earlier. The presence of this neural-network post-filter information indicates that the track includes NNPFC and NNPFA SEI messages. The information may be used by a reader to select which track out of many options with neural-network post-filter information is consumed. aligned(8) class PreselectionGroupBox extends TrackGroupTypeBox('pres')
{
PreselectionProcessingBox preselection_processing; // optional
NNPFInformationBox nnpf_information; // optional
} [0465] In an embodiment, the presence NNPFC and NNPFA SEI messages in preselection indicated in PreselectionProcessingBox may have the following syntax: aligned(8) class PreselectionProcessingBox extends FullBoxfprsp', version=0, flags ){ unsigned int(8) track_order; unsigned int(l) sample_merge_flag; unsigned int(l) nnpf_sei_presence_flag; if (nnpf_sei_presence_flag){ unsigned int(l) nnpf_sei_essential_flag; unsigned int(5) reserved;
} else{ unsigned int(6) reserved;
}
// further attributes and Boxes defining additional processing of
// the track contributing to the preselection
}
[0466] nnpf_sei_presence_flag equal to 1 indicates that the preselection includes tracks storing NNPFC and NNPFA SEI messages.
[0467] nnpf_sei_essential_flag equal to 1 indicates that the processing of tracks storing NNPFC and NNPFA SEI messages is essential.
[0468] nnpf_sei_essential_flag equal to 0 indicates that the processing of tracks storing NNPFC and NNPFA SEI messages is optional.
[0469] In an embodiment, the presence NNPFC and NNPFA SEI messages in preselection indicated in PreselectionTrackGroupEntryBox may have the following syntax aligned(8) class PreselectionTrackGroupEntryBox extends TrackGroupEntryBoxfprse', version, flags)
{ unsigned int(8) numTracks; utf8string preselection_tag; if (flags & 1) { unsigned int(8) selection_priority;
} if (flags & 2) { unsigned int(8) segment_order;
} if (flags & 3) { unsigned int(l) nnpf_sei_presence_flag; if (nnpf_sei_presence_flag){ unsigned int(l) nnpf_sei_essential_flag; unsigned int(5) reserved;
} else{ unsigned int(6) reserved;
}
}
// Boxes describing the preselection
}
[0470] nnpf_sei_presence_flag equal to 1 indicates that the preselection includes tracks storing NNPFC and NNPFA SEI messages.
[0471] nnpf_sei_essential_flag equal to 1 indicates that the processing of tracks storing NNPFC and NNPFA SEI messages is essential.
[0472] nnpf_sei_essential_flag equal to 0 indicates that the processing of tracks storing NNPFC and NNPFA SEI messages is optional.
[0473] Several non-VCL tracks with different post-filters
[0474] In an embodiment, several non-VCL tracks may be present with different post-filters. In an embodiment, the several post-filters may have the same purpose but with different characteristics (e.g., visual quality improvements with different NN-filters). In an alternate embodiment, the several post-filters may have different purposes, with the respective NNPFC and/or NNPFA SEI messages for different purpose stored in different non-VCL tracks. [0475] In an example embodiment, the NNPFC SEI messages for different post-processing filtering purpose or characteristics may comprise and differ in one or more of the following:
• Purpose of the post-filter, for example: o Visual quality improvement o Chroma upsampling from the 4:2:0 chroma format to the 4:2:2 or 4:4:4 chroma format, or from the 4:2:2 chroma format to the 4:4:4 chroma format o Increasing the width or height of the cropped decoded output picture without changing the chroma format o Increasing the width or height of the cropped decoded output picture and upsampling the chroma format
• Formatting of the input tensors that are given as input to the neural network inference
• Formatting of the output tensors that are resulting from the neural network inference
• Characterization of the complexity of the neural network
[0476] Indicating tracks of alternative post-filters with a track group
[0477] When post-filters have the same purpose but different characteristics, the respective NNPFC and/or NNPFA SEI messages could be stored in different non-VCL tracks. The non-VCL tracks may be included in the same track group, such as 'alpp' (alternative post-processing). The track reference 'vvcN' references the ID of the track group, indicating that only one of the tracks in the track group should be selected for processing (e.g., for bitstream reconstruction).
[0478] Indicating tracks of neural-network post-filters within an Entity group
[0479] The neural-network post-filter entity group defines a set of tracks and items which includes NNPFC SEI and the NNPFA SEI messages in the associated bitstream. Furthermore the sneural-network post-filter entity group may specify whether the processing a track or item with a specific entity _id is essential or not.
[0480] Syntax aligned(8) class NNPFEntityGroup extends EntityToGroupBox(’nneg’,0,0)
{ for (i = 0; i < num_entities_in_group; i++) unsigned int(l) nnpf_sei_essential_flag; unsigned int(7) reserved; }
[0481] Semantics
[0482] nnpf_sei_essential_flag equal to 1 indicates that the processing of track or item with the corresponding entity _id storing NNPFC and NNPFA SEI messages is essential.
[0483] nnpf_sei_essential_flag equal to 0 indicates that the processing of track or item with the corresponding entity _id storing NNPFC and NNPFA SEI messages is optional.
[0484] Property descriptor for neural-network post-filter information in DASH MPD
[0485] In an embodiment, a neural- network post-filter information property descriptor is defined to carry syntax elements of NNPFC SEI, but without NNR bitstream. When the property descriptor is used as an essential property, it is required that the neural-network post-filtering is performed. When the property descriptor is used as a supplemental property, it is optional to perform neural-network postfiltering.
[0486] Codecs MIME parameter update with mandatory SEI messages
[0487] In an embodiment, the content of SEI manifest SEI message is included in the codecs MIME parameter. For example, the codecs or some part of it may comprise key-value pairs, where one key (such as seim) may indicate SEI manifest SEI message, and the respective value may for example be a representation of the payload of the SEI manifest SEI message, such as a base64-coded content of the SEI message.
[0488] FIG. 10 is an example apparatus 1000, which may be implemented in hardware, caused to implement mechanisms for integrating neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF), based on the examples described herein. The apparatus 1000 comprises at least one processor 1002, at least one non-transitory memory 1804 including computer program code 1005, wherein the at least one memory 1004 and the computer program code 1005 are configured to, with the at least one processor 1002, cause the apparatus 1000 to implement mechanisms for integrating neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF) 1006, based on the examples described herein. In an embodiment, the at least one neural network or the portion of the at least one neural network may be used at a decoder-side for decoding or reconstructing one or more media items. [0489] The apparatus 1000 optionally includes a display 1008 that may be used to display content during rendering. The apparatus 1000 optionally includes one or more network (NW) interfaces (I/F(s)) 1010. The NW I/F(s) 1010 may be wired and/or wireless and communicate over the Internet/other network(s) via any communication technique. The NW I/F(s) 1010 may comprise one or more transmitters and one or more receivers. The N/W I/F(s) 1010 may comprise standard well-known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder/decoder circuitry(ies) and one or more antennas.
[0490] The apparatus 1000 may be a remote, virtual or cloud apparatus. The apparatus 1000 may be either a coder or a decoder, or both a coder and a decoder. The at least one memory 1004 may be implemented using any suitable data storage technology, such as semiconductor based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The at least one memory 1004 may comprise a database for storing data. The apparatus 1000 need not comprise each of the features mentioned, or may comprise other features as well. The apparatus 1000 may correspond to or be another embodiment of the apparatus 50 shown in FIG. 1 and FIG. 2, any of the apparatuses shown in FIG. 3, or apparatus 700 of FIG. 7. The apparatus 1000 may correspond to or be another embodiment of the apparatuses shown in FIG. 15, including UE 110, RAN node 170, or network element(s) 190.
[0491] FIG. 11 illustrates an example method 1100 for writing a file, in accordance with an embodiment. As shown in block 1006 of FIG.10, the apparatus 1000 includes means, such as the processing circuitry 1002 or the like, for integrating neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF). At 1102, the method 1100 includes storing coded media data in a media track of a file. In an example, the coded media data includes coded video data, the media track includes video track. At 1104, the method 1100 includes storing information in the media track or in a second track associated with the media track. In an example, the information includes storing neural-network post-filter (NNPF) supplemental enhancement information (SEI) and the second track includes a VVC non-VCL track. At 1106, the method 1100 includes indicating, in the file, the presence of the information in the media track or the second track with at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information. In an example, the restricted media sample entry type comprises a restricted video sample entry type. In an example, the NNPFSEI includes at least one of a neural network post filter characteristics (NNPFC )SEI message or a neural network post filter activation (NNPF A) SEI message. [0492] FIG. 12 illustrates an example method 1200 for parsing a file, in accordance with an embodiment. As shown in block 1006 of FIG. 10, the apparatus 1000 includes means, such as the processing circuitry 1002 or the like, for integrating neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF). At 1202, the method 1200 includes reading, from a file, presence of information in a media track or a second track with at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information. At 1204, the method 1200 includes processing the information in the media track or the second track, in response to supporting the information. At 1206, the method 1200 includes decoding a coded media data from the media track. At 1208, the method 1200 filtering the decoded media data with a post-filter defined by the information. In an example, the information includes neural-network prost- filter (NNPF) supplemental enhancement information (SEI), the NNPF SEI includes at least one of a NNPFC SEI message or a NNPFA SEI message, the media track includes a video media track, the second track includes a versatile video coding (VVC) non-video coding layer (non-VCL) track, and/or the media data comprises video data.
[0493] FIG. 13 illustrates another example method 1300 for writing a file, in accordance with an embodiment. At 1302, the method 1300 includes storing coded media data in a media track of a file. At 1304, the method 1300 includes storing information in the media track or in a second track associated with the media track. At 1306, the method 1300 includes indicating, in the file, the presence of the information in the media track or the second track with a sample group for the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
[0494] FIG. 14 illustrates another example method 1400 parsing a file, in accordance with an embodiment. At 1402, the method 1400 includes reading, from a sample group of information in a file, presence of the information in a media track or a second track. At 1404, the method 1400 includes processing the information in the media track or the second track, in response to supporting the information. At 1406, the method 1400 includes decoding a coded media data from the media track; and filtering the decoded media data with a post-filter defined by the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
[0495] Referring to FIG. 15, this figure shows a block diagram of one possible and non-limiting example in which the examples may be practiced. A user equipment (UE) 110, radio access network (RAN) node 170, and network element(s) 190 are illustrated. In the example of FIG. 1, the user equipment (UE) 110 is in wireless communication with a wireless network 100. A UE is a wireless device that can access the wireless network 100. The UE 110 includes one or more processors 120, one or more memories 125, and one or more transceivers 130 interconnected through one or more buses 127. Each of the one or more transceivers 130 includes a receiver, Rx, 132 and a transmitter, Tx, 133. The one or more buses 127 may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. The one or more transceivers 130 are connected to one or more antennas 128. The one or more memories 125 include computer program code 123. The UE 110 includes a module 140, comprising one of or both parts 140-1 and/or 140-2, which may be implemented in a number of ways. The module 140 may be implemented in hardware as module 140-1, such as being implemented as part of the one or more processors 120. The module 140-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the module 140 may be implemented as module 140-2, which is implemented as computer program code 123 and is executed by the one or more processors 120. For instance, the one or more memories 125 and the computer program code 123 may be configured to, with the one or more processors 120, cause the user equipment 110 to perform one or more of the operations as described herein. The UE 110 communicates with RAN node 170 via a wireless link 111.
[0496] The RAN node 170 in this example is a base station that provides access by wireless devices such as the UE 110 to the wireless network 100. The RAN node 170 may be, for example, a base station for 5G, also called New Radio (NR). In 5G, the RAN node 170 may be a NG-RAN node, which is defined as either a gNB or an ng-eNB. A gNB is a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to a 5GC (such as, for example, the network element(s) 190). The ng-eNB is a node providing E-UTRA user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GC. The NG- RAN node may include multiple gNBs, which may also include a central unit (CU) (gNB-CU) 196 and distributed unit(s) (DUs) (gNB-DUs), of which DU 195 is shown. Note that the DU may include or be coupled to and control a radio unit (RU). The gNB-CU is a logical node hosting radio resource control (RRC), SDAP and PDCP protocols of the gNB or RRC and PDCP protocols of the en-gNB that controls the operation of one or more gNB-DUs. The gNB-CU terminates the Fl interface connected with the gNB-DU. The Fl interface is illustrated as reference 198, although reference 198 also illustrates a link between remote elements of the RAN node 170 and centralized elements of the RAN node 170, such as between the gNB-CU 196 and the gNB-DU 195. The gNB-DU is a logical node hosting RLC, MAC and PHY layers of the gNB or en-gNB, and its operation is partly controlled by gNB-CU. One gNB- CU supports one or multiple cells. One cell is supported by only one gNB-DU. The gNB-DU terminates the Fl interface 198 connected with the gNB-CU. Note that the DU 195 is considered to include the transceiver 160, for example, as part of a RU, but some examples of this may have the transceiver 160 as part of a separate RU, for example, under control of and connected to the DU 195. The RAN node 170 may also be an eNB (evolved NodeB) base station, for LTE (long term evolution), or any other suitable base station or node.
[0497] The RAN node 170 includes one or more processors 152, one or more memories 155, one or more network interfaces (N/W I/F(s)) 161, and one or more transceivers 160 interconnected through one or more buses 157. Each of the one or more transceivers 160 includes a receiver, Rx, 162 and a transmitter, Tx, 163. The one or more transceivers 160 are connected to one or more antennas 158. The one or more memories 155 include computer program code 153. The CU 196 may include the processor(s) 152, memories 155, and network interfaces 161. Note that the DU 195 may also include its own memory/memories and processor(s), and/or other hardware, but these are not shown.
[0498] The RAN node 170 includes a module 150, comprising one of or both parts 150-1 and/or 150-2, which may be implemented in a number of ways. The module 150 may be implemented in hardware as module 150-1, such as being implemented as part of the one or more processors 152. The module 150-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the module 150 may be implemented as module 150-2, which is implemented as computer program code 153 and is executed by the one or more processors 152. For instance, the one or more memories 155 and the computer program code 153 are configured to, with the one or more processors 152, cause the RAN node 170 to perform one or more of the operations as described herein. Note that the functionality of the module 150 may be distributed, such as being distributed between the DU 195 and the CU 196, or be implemented solely in the DU 195.
[0499] The one or more network interfaces 161 communicate over a network such as via the links 176 and 131. Two or more gNBs 170 may communicate using, for example, link 176. The link 176 may be wired or wireless or both and may implement, for example, an Xn interface for 5G, an X2 interface for LTE, or other suitable interface for other standards.
[0500] The one or more buses 157 may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, wireless channels, and the like. For example, the one or more transceivers 160 may be implemented as a remote radio head (RRH) 195 for LTE or a distributed unit (DU) 195 for gNB implementation for 5G, with the other elements of the RAN node 170 possibly being physically in a different location from the RRH/DU, and the one or more buses 157 could be implemented in part as, for example, fiber optic cable or other suitable network connection to connect the other elements (for example, a central unit (CU), gNB-CU) of the RAN node 170 to the RRH/DU 195. Reference 198 also indicates those suitable network link(s).
[0501] It is noted that description herein indicates that ‘cells’ perform functions, but it should be clear that equipment which forms the cell may perform the functions. The cell makes up part of a base station. That is, there can be multiple cells per base station. For example, there could be three cells for a single carrier frequency and associated bandwidth, each cell covering one-third of a 360 degree area so that the single base station’s coverage area covers an approximate oval or circle. Furthermore, each cell can correspond to a single carrier and a base station may use multiple carriers. So when there are three 120 degree cells per carrier and two carriers, then the base station has a total of 6 cells.
[0502] The wireless network 100 may include a network element or elements 190 that may include core network functionality, and which provides connectivity via a link or links 181 with a further network, such as a telephone network and/or a data communications network (for example, the Internet). Such core network functionality for 5G may include access and mobility management function(s) (AMF(S)) and/or user plane functions (UPF(s)) and/or session management function(s) (SMF(s)). Such core network functionality for LTE may include MME (Mobility Management Entity )/SGW (Serving Gateway) functionality. These are merely example functions that may be supported by the network element(s) 190, and note that both 5G and LTE functions might be supported. The RAN node 170 is coupled via a link 131 to the network element 190. The link 131 may be implemented as, for example, an NG interface for 5G, or an SI interface for LTE, or other suitable interface for other standards. The network element 190 includes one or more processors 175, one or more memories 171, and one or more network interfaces (N/W I/F(s)) 180, interconnected through one or more buses 185. The one or more memories 171 include computer program code 173. The one or more memories 171 and the computer program code 173 are configured to, with the one or more processors 175, cause the network element 190 to perform one or more operations.
[0503] The wireless network 100 may implement network virtualization, which is the process of combining hardware and software network resources and network functionality into a single, softwarebased administrative entity, a virtual network. Network virtualization involves platform virtualization, often combined with resource virtualization. Network virtualization is categorized as either external, combining many networks, or parts of networks, into a virtual unit, or internal, providing network-like functionality to software containers on a single system. Note that the virtualized entities that result from the network virtualization are still implemented, at some level, using hardware such as processors 152 or 175 and memories 155 and 171, and also such virtualized entities create technical effects. [0504] The computer readable memories 125, 155, and 171 may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The computer readable memories 125, 155, and 171 may be means for performing storage functions. The processors 120, 152, and 175 may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on a multi-core processor architecture, as non-limiting examples. The processors 120, 152, and 175 may be means for performing functions, such as controlling the UE 110, RAN node 170, network element(s) 190, and other functions as described herein.
[0505] In general, the various embodiments of the user equipment 110 may include, but are not limited to, cellular telephones such as smart phones, tablets, personal digital assistants (PDAs) having wireless communication capabilities, portable computers having wireless communication capabilities, image capture devices such as digital cameras having wireless communication capabilities, gaming devices having wireless communication capabilities, music storage and playback appliances having wireless communication capabilities, Internet appliances permitting wireless Internet access and browsing, tablets with wireless communication capabilities, as well as portable units or terminals that incorporate combinations of such functions.
[0506] One or more of modules 140-1, 140-2, 150-1, and 150-2 may be configured to integrate neural -network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF). Computer program code 173 may also be configured to integrate neural-network post-filter supplemental enhancement information (SEI) with ISO base media file format (ISOBMFF).
[0507] As described above, FIGs. 11 to 14 include flowcharts of an apparatus (e.g. 50, 100, 602, 604, 700, or 1000), method, and computer program product according to certain example embodiments. It will be understood that each block of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by various means, such as hardware, firmware, processor, circuitry, and/or other devices associated with execution of software including one or more computer program instructions. For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instructions which embody the procedures described above may be stored by a memory (e.g. 58, 125, 704, or 1004) of an apparatus employing an embodiment of the present invention and executed by processing circuitry (e.g. 56, 120, 702, or 1002) of the apparatus. As will be appreciated, any such computer program instructions may be loaded onto a computer or other programmable apparatus (e.g., hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture, the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.
[0508] A computer program product is therefore defined in those instances in which the computer program instructions, such as computer-readable program code portions, are stored by at least one non- transitory computer-readable storage medium with the computer program instructions, such as the computer-readable program code portions, being configured, upon execution, to perform the functions described above, such as in conjunction with the flowchart(s) of FIGs. 11 to 14. In other embodiments, the computer program instructions, such as the computer-readable program code portions, need not be stored or otherwise embodied by a non-transitory computer-readable storage medium, but may, instead, be embodied by a transitory medium with the computer program instructions, such as the computer- readable program code portions, still being configured, upon execution, to perform the functions described above.
[0509] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions for performing the specified functions. It will also be understood that one or more blocks of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by special purpose hardware-based computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.
[0510] In some embodiments, certain ones of the operations above may be modified or further amplified. Furthermore, in some embodiments, additional optional operations may be included. Modifications, additions, or amplifications to the operations above may be performed in any order and in any combination.
[0511] In the above, some embodiments have been described with reference to NNPFC SEI message(s) and/or NNPFA SEI message(s). It needs to be understood that embodiments can be similarly realized with any SEI messages of similar nature. For example, embodiments may be realized with postfilter characteristics and/or activation SEI message(s) where post-filters are not based on neural networks. In another example, there may be more than one SEI message that provides properties for post-filtering. In yet another example, the NNPFA SEI message may have a persistence over multiple pictures, in which case it is not inserted with each coded picture in the reconstructed bitstream but rather only with the first coded picture, in decoding order, within a sequence of pictures that the NNPFA SEI message pertains to.
[0512] In the above, some embodiments have been described with reference to NNPFC SEI message(s) and/or NNPFA SEI message(s), but it needs to be understood that they generally apply to any type of SEI messages.
[0513] In the above, some example embodiments have been described with reference to an SEI message or an SEI NAL unit. It needs to be understood, however, that embodiments can be similarly realized with any similar structures or data units, such as metadata OBUs. Where example embodiments have been described with SEI messages included in a structure, any independently parsable structures could likewise be used in embodiments. Specific SEI NAL unit and a SEI message syntax structures have been presented in example embodiments, but it needs to be understood that embodiments generally apply to any syntax structures with a similar intent as SEI NAL units and/or SEI messages.
[0514] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and/or computer program may reside at the encoder for generating the bitstream and/or at the decoder for decoding the bitstream.
[0515] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and/or computer program for generating the bitstream to be decoded by the decoder.
[0516] In the above, some embodiments have been described with the help of syntax of a file format, such as ISOBMFF. It needs to be understood, however, that the corresponding structure and/or computer program may reside at a file writer for generating a file and/or at a file reader for parsing or interpreting a file. [0517] In the above, where example embodiments have been described with reference to a file writer, it needs to be understood that the resulting file and the file reader have corresponding elements in them. Likewise, where example embodiments have been described with reference to a file reader, it needs to be understood that the file writer has structure and/or computer program for generating the file to be parsed or interpreted by the file reader.
[0518] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, although the foregoing descriptions and the associated drawings describe example embodiments in the context of certain example combinations of elements and/or functions, it should be appreciated that different combinations of elements and/or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, different combinations of elements and/or functions than those explicitly described above are also contemplated as may be set forth in some of the appended claims. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0519] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.
[0520] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not only computers having different architectures such as single/multi-processor architectures and sequential (Von Neumann)/parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device, and the like. [0521] As used herein, the term ‘circuitry’ or ‘’circuit’ may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and/or digital circuitry, and (b) combinations of circuits and software (and/or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s)/software including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even when the software or firmware is not physically present. This description of ‘circuitry’ or ‘circuit’ applies to uses of this term in this application. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and/or firmware. The term ‘circuitry’ or ‘circuit’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.

Claims

CLAIMS What is claimed is:
1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: store coded media data in a media track of a file; store information in the media track or in a second track associated with the media track; and indicate, in the file, the presence of the information in the media track or the second track with a sample group for the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
2. The apparatus of claim 1, wherein the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPF A) SEI message.
3. The apparatus of any of the previous claims, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track.
4. The apparatus of any of the previous claims, wherein the apparatus is further caused to define a sample group description entry to carry syntax elements of the NNPFC SEI message.
5. The apparatus of any of the previous claims, wherein the apparatus is further caused to define a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘ 1 ’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box comprises an NNPFC identifier equal to a filter identifier.
6. The apparatus of claim 2, wherein the NNPFA SEI message comprises an NNPFA identifier syntax element to identify a post-processing filter that the NNPFA SEI message is associated with.
7. The apparatus of claim 1 or 2, wherein an applicable post-processing filter with an NNPFC identifier equal to a NNPFA identifier is used to filter an image/picture comprising the NNPFA SEI message.
8. The apparatus of any of the previous claims, wherein the apparatus is further caused to indicate the sample group as an essential sample group and the player is required to process the information when the sample group is indicated as the essential sample group.
9. The apparatus of claim 1, wherein the apparatus is further caused to include a content of SEI manifest SEI message in a codecs multipurpose internet mail extension (MIME) parameter, wherein the codecs or a part of the codecs comprises key -value pairs, wherein a key of the key-value pair indicates SEI manifest SEI message, and a respective value comprises representation of a payload of the SEI manifest SEI message.
10. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: read, from a sample group for information in a file, presence of the information in a media track or a second track; process the information in the media track or the second track, in response to supporting the information; decode a coded media data from the media track; and filter the decoded media data with a post-filter defined by the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
11. The apparatus of claim 10, wherein: the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPF A) SEI message.
12. The apparatus of claim 10 or 11, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track.
13. The apparatus of any of the claims 10 to 12, wherein the apparatus is caused to process the sample group, wherein the first sample group description entry comprises the NNPFC SEI message, and the apparatus is caused to read a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘1’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the sample group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box include an NNPFC identifier equal to filter identifier.
14. The apparatus of claim 13, wherein to process the first sample group, the apparatus is caused to: ignore and skip the VVC non-VCL track in a bitstream reconstruction, when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus does not recognize the first sample group; or perform following insertion of prefix SEI NAL units when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus recognizes the first sample group, and wherein following insertion is optional when the first sample group is a non-essential group; when a sample is mapped to at least one sample group description entry of the sample group and the sample is either a sync sample or the first sample of a sequence of samples associated with the same sample entry, the sample implicitly comprises a prefix SEI NAL unit for each layer comprised in the media track and each filter identifier value mapped to the sample, and wherein the prefix SEI NAL unit comprises the NNPFC SEI message from the sample group description entry with a filter update flag equal to 0, followed by the NNPFC SEI message from the sample group description entry with filter update flag equal to 1, when present.
15. The apparatus of any of the claims 10 to 14, wherein the media track or the second track comprises at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information.
16. An method comprising: storing coded media data in a media track of a file; storing information in the media track or in a second track associated with the media track; and indicating, in the file, the presence of the information in the media track or the second track with a sample group for the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
17. The method of claim 16, wherein the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPF A) SEI message.
18. The method of any of the claims 16 or 17, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track.
19. The method of any of the claims 16 to 18 further comprising defining a sample group description entry to carry syntax elements of the NNPFC SEI message.
20. The method of any of the claims 16 to 19 further comprising defining a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘ 1 ’ indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the group description entry referenced by the mapping box comprises the NNPFC SEI message that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box comprises an NNPFC identifier equal to a filter identifier.
21. The method of claim 17, wherein the NNPFA SEI message comprises an NNPFA identifier syntax element to identify a post-processing filter that the NNPFA SEI message is associated with.
22. The method of claim 16 or 17, wherein an applicable post-processing filter with an NNPFC identifier equal to a NNPFA identifier is used to filter an image/picture comprising the NNPFA SEI message.
23. The method of any of the claims 16 to 22 further comprising indicating the sample group as an essential sample group and the player is required to process the information when the sample group is indicated as the essential sample group.
24. The method of claim 16 further comprising including a content of SEI manifest SEI message in a codecs multipurpose internet mail extension (MIME) parameter, wherein the codecs or a part of the codecs comprises key-value pairs, wherein a key of the key-value pair indicates SEI manifest SEI message, and a respective value comprises representation of a payload of the SEI manifest SEI message.
25. An method comprising: reading, from a sample group of information in a file, presence of the information in a media track or a second track; processing the information in the media track or the second track, in response to supporting the information; decoding a coded media data from the media track; and filtering the decoded media data with a post-filter defined by the information, wherein the information comprises a neural network post filter (NNPF) supplemental enhancement information (SEI).
26. The method of claim 25, wherein: the neural network post filter (NNPF) supplemental enhancement information (SEI) comprises at least one of a neural network post filter characteristics (NNPFC) SEI message or a neural network post filter activation (NNPFA) SEI message.
27. The method of any of the claims 25 or 26, wherein, the coded media data comprise coded video data, media track comprises a video track, and the second video track comprises a versatile video coding (VVC) non-video coding layer (non-VCL) track.
28. The method of any of the claims 25 to 27 further comprising: processing the sample group, wherein the first sample group description entry comprises the NNPFC SEI message; and read a first grouping type parameter for a sample-to-group box of the sample group, wherein the first grouping type parameter comprises: a flag, wherein the flag equal to ‘T indicates that the sample group description entry referenced by a first mapping box comprises the NNPFC SEI message that provides an update on top of a post-processing filter, and wherein the flag equal to ‘0’ indicates that the sample group description entry referenced by the mapping box comprises the NNPFC SEI message specifies that specifies a base post processing filter; and a filter identifier to indicate that the sample group description entry referenced by the mapping box include an NNPFC identifier equal to filter identifier.
29. The method of claim 28, wherein to process the first sample group, the method further comprises: ignoring and skip the VVC non-VCL track in a bitstream reconstruction, when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and an apparatus does not recognize the first sample group; or performing following insertion of prefix SEI NAL units when the first sample group is an essential sample group that is present in the VVC non-VCL NAL unit and the apparatus recognizes the first sample group, and wherein following insertion is optional when the first sample group is a non-essential group; when a sample is mapped to at least one sample group description entry of the sample group and the sample is either a sync sample or the first sample of a sequence of samples associated with the same sample entry, the sample implicitly comprises a prefix SEI NAL unit for each layer comprised in the media track and each filter identifier value mapped to the sample, and wherein the prefix SEI NAL unit comprises the NNPFC SEI message from the sample group description entry with a filter update flag equal to 0, followed by the NNPFC SEI message from the sample group description entry with filter update flag equal to 1, when present.
30. The method of any of the claims 25 to 29, wherein the media track or the second track comprises at least one of a scheme type of a restricted media sample entry type or an essential sample group for the information.
31. A computer -readable medium encoded with instructions that, when executed by a computer, causing an apparatus to perform a method according to any of the claims 16 to 24.
32. The computer-readable medium of claim 31, wherein the computer-readable medium comprises a non-transitory computer-readable medium.
33. An apparatus comprising means for performing the methods as claimed in any of the claims 16 to 24.
34. A computer -readable medium encoded with instructions that, when executed by a computer, causing an apparatus to perform a method according to any of the claims 25 to 30.
35. The computer -readable medium of claim 34, wherein the computer-readable medium comprises a non-transitory computer-readable medium.
36. An apparatus comprising means for performing the methods as claimed in any of the claims 25 to 30.
EP23793066.4A 2022-10-14 2023-10-13 Apparatus and method for integrating neural-network post-filter supplemental enhancement information with iso base media file format Pending EP4602807A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202263379540P 2022-10-14 2022-10-14
PCT/IB2023/060365 WO2024079718A1 (en) 2022-10-14 2023-10-13 Apparatus and method for integrating neural-network post-filter supplemental enhancement information with iso base media file format

Publications (1)

Publication Number Publication Date
EP4602807A1 true EP4602807A1 (en) 2025-08-20

Family

ID=88504706

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23793066.4A Pending EP4602807A1 (en) 2022-10-14 2023-10-13 Apparatus and method for integrating neural-network post-filter supplemental enhancement information with iso base media file format

Country Status (2)

Country Link
EP (1) EP4602807A1 (en)
WO (1) WO2024079718A1 (en)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20240015332A1 (en) * 2022-07-08 2024-01-11 Qualcomm Incorporated Supplemental enhancement information (sei) manifest indication

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2022079545A1 (en) 2020-10-13 2022-04-21 Nokia Technologies Oy Carriage and signaling of neural network representations

Also Published As

Publication number Publication date
WO2024079718A1 (en) 2024-04-18

Similar Documents

Publication Publication Date Title
US12113974B2 (en) High-level syntax for signaling neural networks within a media bitstream
US12036036B2 (en) High-level syntax for signaling neural networks within a media bitstream
RU2746934C9 (en) Interlevel prediction for scalable encoding and decoding of video information
KR102089457B1 (en) Apparatus, method and computer program for video coding and decoding
CN107431810B (en) Apparatus, method and computer program for image encoding and decoding
US12219204B2 (en) Carriage and signaling of neural network representations
US20250274611A1 (en) Method For Signaling Media Usability For Neural Networks
US20250254403A1 (en) Method and apparatus for encoding, decoding, or displaying picture-in-picture
US12068007B2 (en) Method, apparatus and computer program product for signaling information of a media track
EP4602807A1 (en) Apparatus and method for integrating neural-network post-filter supplemental enhancement information with iso base media file format
US12495152B2 (en) Method and apparatus for encoding, decoding, or progressive rendering of image
US20240397128A1 (en) Method, apparatus and computer program product for implementing mechanisms for carriage of renderable text
US12489907B2 (en) Method and apparatus for signaling of regions and region masks in image file format
US12621519B2 (en) Carriage and signaling of neural network representations
US20250317618A1 (en) Method and apparatus for enabling backward compatible multilayer bitstream in single layer tracks
WO2025012831A1 (en) Carriage of neural-network post-filter
WO2023194816A1 (en) Method and apparatus for tracking group entry information
WO2025078976A1 (en) A method an apparatus and a computer program for encapsulating and streaming attenuation maps for green metadata
WO2025153959A1 (en) Slim mode for image file format

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250514

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)