EP4684531A1 - A method, an apparatus and a computer program product for video coding - Google Patents

A method, an apparatus and a computer program product for video coding

Info

Publication number
EP4684531A1
EP4684531A1 EP24774313.1A EP24774313A EP4684531A1 EP 4684531 A1 EP4684531 A1 EP 4684531A1 EP 24774313 A EP24774313 A EP 24774313A EP 4684531 A1 EP4684531 A1 EP 4684531A1
Authority
EP
European Patent Office
Prior art keywords
media stream
media
header extension
real time
transport packet
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24774313.1A
Other languages
German (de)
French (fr)
Inventor
Lukasz Kondrad
Igor Danilo Diego Curcio
Sujeet Shyamsundar Mate
Serhan GÜL
Saba AHSAN
Kashyap KAMMACHI SREEDHAR
Lauri Aleksi ILOLA
Emre Baris Aksu
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Nokia Technologies Oy
Original Assignee
Nokia Technologies Oy
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nokia Technologies Oy filed Critical Nokia Technologies Oy
Publication of EP4684531A1 publication Critical patent/EP4684531A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/60Network structure or processes for video distribution between server and client or between remote clients; Control signalling between clients, server and network components; Transmission of management data between server and client, e.g. sending from server to client commands for recording incoming content stream; Communication details between server and client 
    • H04N21/63Control signaling related to video distribution between client, server and network components; Network processes for video distribution between server and clients or between remote clients, e.g. transmitting basic layer and enhancement layers over different transmission paths, setting up a peer-to-peer communication via Internet between remote STB's; Communication protocols; Addressing
    • H04N21/643Communication protocols
    • H04N21/6437Real-time Transport Protocol [RTP]
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L65/00Network arrangements, protocols or services for supporting real-time applications in data packet communication
    • H04L65/60Network streaming of media packets
    • H04L65/65Network streaming protocols, e.g. real-time transport protocol [RTP] or real-time control protocol [RTCP]
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L65/00Network arrangements, protocols or services for supporting real-time applications in data packet communication
    • H04L65/60Network streaming of media packets
    • H04L65/75Media network packet handling
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/81Monomedia components thereof
    • H04N21/816Monomedia components thereof involving special video data, e.g 3D video
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L65/00Network arrangements, protocols or services for supporting real-time applications in data packet communication
    • H04L65/40Support for services or applications
    • H04L65/403Arrangements for multi-party communication, e.g. for conferences
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L65/00Network arrangements, protocols or services for supporting real-time applications in data packet communication
    • H04L65/80Responding to QoS
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L69/00Network arrangements, protocols or services independent of the application payload and not provided for in the other groups of this subclass
    • H04L69/22Parsing or analysis of headers
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L69/00Network arrangements, protocols or services independent of the application payload and not provided for in the other groups of this subclass
    • H04L69/24Negotiation of communication capabilities
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof
    • H04N13/10Processing, recording or transmission of stereoscopic or multi-view image signals
    • H04N13/106Processing image signals
    • H04N13/172Processing image signals image signals comprising non-image signal components, e.g. headers or format information
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof
    • H04N13/10Processing, recording or transmission of stereoscopic or multi-view image signals
    • H04N13/194Transmission of image signals

Definitions

  • the present solution generally relates to coding of volumetric video.
  • Volumetric video data represents a three-dimensional (3D) scene or object, and can be used as input for AR (Augmented Reality), VR (Virtual Reality), and MR (Mixed Reality) applications.
  • AR Augmented Reality
  • VR Virtual Reality
  • MR Magnetic Reality
  • Such data describes geometry (Shape, size, position in 3D space) and respective attributes (e.g., color, opacity, reflectance, . ..), and any possible temporal transformations of the geometry and attributes at given time instances (like frames in 2D video).
  • Volumetric video can be generated from 3D models, also referred to as volumetric visual objects, i.e., CGI (Computer Generated Imagery), or captured from real-world scenes using a variety of capture solutions, e.g., multi-camera, laser scan, combination of video and dedicated depth sensors, and more. Also, a combination of CGI and real-world data is possible.
  • CGI Computer Generated Imagery
  • V3C visual volumetric video-based coding
  • video components containing occupancy, geometry or attribute information may be encapsulated into Real-time Transport Protocol (RTP) video streams.
  • RTP Real-time Transport Protocol
  • the problem presented above may be solved by introducing a new feature to a session negotiation process, e.g. in using the session description protocol (e.g., SDP) for offer-answer negotiation.
  • the new feature would consist of a signaling mechanism during session negotiation that informs a client that a RTP header extension information provided by one media stream may be utilized for processing another media stream.
  • a new signalling is provided by additional attributes in the RTP header extension uniform resource identifier (URI) / identifier (ID) to media mapping.
  • URI uniform resource identifier
  • ID identifier
  • such signalling would be implicitly provided by the extension definition, i.e. URI or ID, that a given RTP header extension is applicable to all media present in the SDP group :bundle.
  • such signaling could be provided in the RTP header extension itself utilizing the media id MID.
  • the RTP header extension is attached to multiple streams at different frames. This signalling would indicate that the header extension is flexibly present in different streams.
  • the session description document contains information indicating that the RTP header extension of the first media stream can be utilized for the processing of the second media stream.
  • an apparatus comprising means for forming a session description containing information about at least a first media stream and a second media stream; means for receiving at least the first media stream and the second media stream; means for associating the first media stream a real time transport packet header extension; and means for providing reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
  • a method comprising forming a session description containing information about at least a first media stream and a second media stream; receiving at least the first media stream and the second media stream; associating the first media stream a real time transport packet header extension; and providing reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
  • an apparatus comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following: forming a session description containing information about at least a first media stream and a second media stream; receiving at least the first media stream and the second media stream; associating the first media stream a real time transport packet header extension; and providing reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
  • computer program product comprising computer program code configured to, when executed on at least one processor, cause an apparatus or a system to form a session description containing information about at least a first media stream and a second media stream; receive at least the first media stream and the second media stream; associate the first media stream a real time transport packet header extension; and provide reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
  • the second media stream is left without a real time transport packet header extension.
  • the reuse information is provided in the real time transport packet header extension.
  • identification of the second media stream is provided in the reuse information.
  • the real time transport packet header extension is attached to multiple streams at different frames, wherein the reuse information indicates that the real time transport packet header extension is flexibly present in different streams.
  • the reuse information is included in extension attributes in the real time transport packet header extension.
  • the first media stream and the second media stream are streams of different components of a volumetric video.
  • the computer program product is embodied on a non-transitory computer readable medium.
  • an apparatus comprising means for receiving a session description containing information about at least a first media stream and a second media stream; means for determining that the first media stream is associated with a real time transport packet header extension; means for parsing a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and means for using the real time transport packet header extension of the first media stream for processing the second media stream.
  • a method comprising receiving a session description containing information about at least a first media stream and a second media stream; determining that the first media stream is associated with a real time transport packet header extension; parsing a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and using the real time transport packet header extension of the first media stream for processing the second media stream.
  • an apparatus comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following: receiving a session description containing information about at least a first media stream and a second media stream; determining that the first media stream is associated with a real time transport packet header extension; parsing a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and using the real time transport packet header extension of the first media stream for processing the second media stream.
  • computer program product comprising computer program code configured to, when executed on at least one processor, cause an apparatus or a system to receive a session description containing information about at least a first media stream and a second media stream; determine that the first media stream is associated with a real time transport packet header extension; parse a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and use the real time transport packet header extension of the first media stream for processing the second media stream.
  • the second media stream is left without a real time transport packet header extension.
  • the reuse information is parsed from the real time transport packet header extension.
  • identification of the second media stream is parsed from the reuse information.
  • the real time transport packet header extension is attached to multiple streams at different frames, wherein the reuse information indicates that the real time transport packet header extension is flexibly present in different streams.
  • the reuse information is received in extension attributes in the real time transport packet header extension.
  • Some embodiments may allow to avoid duplication of information in a transmission system. Furthermore, some embodiments enable over-riding legacy header extension information without the need for deep inspection and manipulation of the header extension.
  • FIG. 1 shows an example of a session initiation between a server and a client
  • FIG. 2 illustrates a scenario in which a server receives media from another entity such as another client device
  • FIG. 3 illustrates a possible architecture on how a V3C bitstream can be delivered over network using RTP protocols
  • Fig. 4 shows an example of an extension header
  • FIG. 5 is a flowchart illustrating a method according to an embodiment
  • FIG. 6 is a flowchart illustrating a method according to another embodiment
  • FIG. 7 shows an apparatus according to an embodiment
  • FIG. 8 shows an apparatus according to another embodiment
  • Fig. 9 shows an apparatus according to an embodiment.
  • the present embodiments are in particularly targeted to a solution for reusing a header extension delivered by one media by other media in a bundle.
  • volumetric frame can be represented as a point cloud.
  • a point cloud is a set of unstructured points in 3D space, where each point is characterized by its position in a 3D coordinate system (e.g., Euclidean), and some corresponding attributes (e.g., color information provided as RGBA value, or normal vectors).
  • a volumetric frame can be represented as images, with or without depth, captured from multiple viewpoints in 3D space.
  • the volumetric video can be represented by one or more view frames (where a view is a projection of a volumetric scene on to a plane (the camera plane) using a real or virtual camera with known/ computed extrinsic and intrinsic).
  • Each view may be represented by a number of components (e.g., geometry, color, transparency, and occupancy picture), which may be part of the geometry picture or represented separately.
  • a volumetric frame can be represented as a mesh.
  • Mesh is a collection of points, called vertices, and connectivity information between vertices, called edges. Vertices along with edges form faces. The combination of vertices, edges and faces can uniquely approximate shapes of objects.
  • a volumetric frame can provide viewers the ability to navigate a scene with six degrees of freedom, i.e., both translational and rotational movement of their viewing pose (which includes yaw, pitch, and roll).
  • the data to be coded for a volumetric frame can also be significant, as a volumetric frame can contain many numbers of objects, and the positioning and movement of these objects in the scene can result in many dis-occluded regions.
  • the interaction of the light and materials in objects and surfaces in a volumetric frame can generate complex light fields that can produce texture variations for even a slight change of pose.
  • a sequence of volumetric frames is a volumetric video. Due to large amount of information, storage and transmission of a volumetric video requires compression.
  • a way to compress a volumetric frame can be to project the 3D geometry and related attributes into a collection of 2D images along with additional associated metadata.
  • the projected 2D images can then be coded using 2D video and image coding technologies, for example ISO/IEC 14496-10 (H.264/AVC) and ISO/IEC 23008-2 (H.265/HEVC).
  • the metadata can be coded with technologies specified in specification such as ISO/IEC 23090-5.
  • the coded images and the associated metadata can be stored or transmitted to a client that can decode and render the 3D volumetric frame.
  • a Real time Transport Protocol is intended for an end-to-end, real-time transport or streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery.
  • RTP allows data transport to multiple destinations through IP multicast or to a specific destination through IP unicast.
  • the majority of the RTP implementations are built on the User Datagram Protocol (UDP).
  • UDP User Datagram Protocol
  • Other transport protocols may also be utilized.
  • RTP is used in together with other protocols such as H.323 and Real Time Streaming Protocol RTSP.
  • RTP Resource Transfer Protocol
  • RTCP RTP Control Protocol
  • RTP sessions may be initiated between client and server using a signalling protocol, such as H.323, the Session Initiation Protocol (SIP), or RTSP. These protocols may use the Session Description Protocol (RFC 8866) to specify the parameters for the sessions.
  • a signalling protocol such as H.323, the Session Initiation Protocol (SIP), or RTSP.
  • SIP Session Initiation Protocol
  • RTSP Real-Time Transport Protocol
  • RRC 8866 Session Description Protocol
  • RTP is designed to carry a multitude of multimedia formats, which permits the development of new formats without revising the RTP standard.
  • the information required by a specific application of the protocol is not included in the generic RTP header.
  • an RTP profile may be defined.
  • an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may require a profile and payload format specifications.
  • the profile defines the codecs used to encode the payload data and their mapping to payload format codecs in the protocol field Payload Type (PT) of the RTP header.
  • PT Payload Type
  • RTP profile for audio and video conferences with minimal control is defined in RFC 3551.
  • the profile defines a set of static payload type assignments, and a dynamic mechanism for mapping between a payload format, and a PT value using Session Description Protocol (SDP).
  • SDP Session Description Protocol
  • the latter mechanism is used for newer video codec such as RTP payload format for H.264 Video defined in RFC 6184 or RTP Payload Format for High Efficiency Video Coding (HEVC) defined in RFC 7798.
  • An RTP session is established for each multimedia stream. Audio and video streams may use separate RTP sessions, enabling a receiver to selectively receive components of a particular stream.
  • the RTP specification recommends even port number for RTP, and the use of the next odd port number of the associated RTCP session. A single port can be used for RTP and RTCP in applications that multiplex the protocols.
  • Each RTP stream consists of RTP packets, which in turn consist of RTP header and payload pairs.
  • RTP packets are created at the application layer and handed to the transport layer for delivery. Each unit of RTP media data created by an application begins with the RTP packet header.
  • the RTP header has a minimum size of 12 bytes. After the header, optional header extensions may be present. This is followed by the RTP payload, the format of which is determined by the particular class of application.
  • the fields in the header are as follows:
  • P (Padding) (1 bit) Used to indicate if there are extra padding bytes at the end of the RTP packet.
  • Extension header (1 bit) Indicates the presence of an extension header between the header and payload data.
  • the extension header is application or profile specific.
  • CSRC count (4 bits) Contains the number of CSRC identifiers that follow the SSRC • M (Marker) : (1 bit) Signalling used at the application level in a profile-specific manner. If it is set, it means that the current data has some special relevance for the application.
  • PT Payload type: (7 bits) Indicates the format of the payload and thus determines its interpretation by the application.
  • Sequence number (16 bits) The sequence number is incremented for each RTP data packet sent and is to be used by the receiver to detect packet loss and to accommodate out-of-order delivery.
  • Timestamp (32 bits) Used by the receiver to play back the received samples at appropriate time and interval. When several media streams are present, the timestamps may be independent in each stream. The granularity of the timing is application specific. For example, video stream may use a 90 kHz clock. The clock granularity is one of the details that is specified in the RTP profile for an application.
  • Synchronization source identifier uniquely identifies the source of the stream. The synchronization sources within the same RTP session will be unique.
  • Header extension (optional, presence indicated by Extension field)
  • the first 32-bit word contains a profile-specific identifier (16 bits) and a length specifier (16 bits) that indicates the length of the extension in 32-bit units, excluding the 32 bits of the extension header.
  • the extension header data is shown in Fig. 4.
  • the RTP specification RFC3550 provides a capability to extend the RTP header.
  • Section 5.3.1 of RFC3550 defines the header extension format and rules for its use. Header extensions may carry metadata in addition to the usual RTP header information, provided the RTP layer can function if that metadata is missing.
  • the RTP specification RFC3550 states that RTP "is designed so that the header extension may be ignored by other interoperating implementations that have not been extended.” The intent of this restriction is that RTP header extensions (RTP HE) must not be used to extend RTP itself in a manner that is backward incompatible with non-extended implementations.
  • RTP HE RTP header extensions
  • the general mechanism for HE specification, RFC8285 provides the option to use a small number of extensions in each RTP packet, where the domain of possible extensions is large and registration is decentralized.
  • SDP Session Description Protocol
  • SDP is used as an example of a session specific file format.
  • SDP is a format for describing multimedia communication sessions for the purposes of announcement and invitation. Its predominant use is in support of conversational and streaming media applications. SDP does not deliver any media streams itself, but is used between endpoints for negotiation of network metrics, media types, bandwidth requirements, and other associated properties. The set of properties and parameters is called a session profile. SDP is extensible for the support of new media types and formats.
  • the Session Description Protocol describes a session as a group of fields in a text-based format, one field per line.
  • Session descriptions consist of three sections: session, timing, and media descriptions. Each description may contain multiple timing and media descriptions. Names are only unique within the associated syntactic construct.
  • Examples of attributes defined in RFC8866 are “rtpmap” and “fimtp”.
  • the media types are “audio/L8” and “audio/L16”.
  • Codecspecific parameters may be added in other attributes, for example, "fimtp".
  • "fimtp" attribute allows parameters that are specific to a particular format to be conveyed in a way that SDP does not have to understand them.
  • the format can be one of the formats specified for the media. Format-specific parameters, semicolon separated, may be any set of parameters required to be conveyed by SDP and given unchanged to the media tool that will use this format. At most one instance of this attribute is allowed for each format.
  • RFC7798 defines the following sprop-vps, sprop-sps, sprop-pps, profile-space, profile-id, tier-flag, level-id, interop-constraints, profile-compatibility- indicator, sprop-sub-layer-id, recv-sub-layer-id, max-recv-level-id, tx-mode, max-lsr, max- Ips, max-cpb, max-dpb, max-br, max-tr, max-tc, max-fps, sprop-max-don-diff, sprop- depack-buf-nalus, sprop-depack-buf-bytes, depack-buf-cap, sprop-segmentation-id, sprop- spatial-segmentation-idc, dec-parallel-cap, and include-dph.
  • ‘group” and “mid” attributes defined in RFC 5888 allows to group "m" lines in SDP for different purposes.
  • An example can be for lip synchronization or for receiving a media flow consisting of several media streams on different transport addresses.
  • each "m" line is identified by a token, which is carried in a "mid” attribute below the “m” line.
  • the session description carries session-level “group” attributes that group different "m” lines (identified by their tokens) using different group semantics.
  • the semantics of a group describe the purpose for which the "m” lines are grouped.
  • the "group” line indicates that the "m” lines identified by tokens 1 and 2 (the audio and the video "m” lines, respectively) are grouped for the purpose of lip synchronization (LS).
  • RFC5888 defines two semantics for group Lip Synchronization (LS), as used in the example above, and Flow Identification (FID).
  • RFC5583 defines another grouping type Decoding Dependency (DDP).
  • RFC8843 defines another grouping type BUNDLE, which among other is utilized when multiple types of media are sent in a single RTP session as described in RFC8860.
  • Two media-level and one session-level attributes are used in a mechanism for providing alternative SDP lines.
  • One or more SDP lines at media level can be replaced, if desired, by alternatives.
  • the mechanism is backwards compatible in the way that a receiver that does not support the attributes may get the default configuration.
  • the different alternatives can be grouped using different attributes that can be specified hierarchically with a top and a lower level.
  • 3 GPP Release 6 supports grouping based on bit-rate, according to the SDP bandwidth modifiers AS [RFC8866] and TIAS [RFC3890], and language.
  • the SDP attributes are:
  • Language and bit-rate are two defined grouping types.
  • ⁇ value> is the local identifier (ID) of this extension and is an integer in the valid range (0 is reserved for padding in both forms, and 15 is reserved in the one-byte header form, as noted above).
  • ⁇ URI> is a URI, as above
  • the SDP attribute extmap is a SPECIAL category SDP attribute as defined in RFC 8859.
  • the text in the specification defining the attribute must be consulted for further handling when multiplexed. Therefore, multiplexing considerations for RTP header extensions need to be defined by the specification of every RTP header extension.
  • Remote Rendering offloads rendering process of 3D content to a cloud and streams it to devices in real time. Due to that an end user may be able to use interactive, high-quality 3D models with every detail intact and no compromise in quality on a devices with limited capabilities.
  • a client may send continuous metadata information to a server based on which appropriate media is provided to the client.
  • a client may send its pose in its local virtual environment, based on the pose the server renders the 3D scene into 2D views, and subsequently encodes the 2D views into a video which can be streamed to the client.
  • gaming platforms utilize typically WebRTC data channels (SCTP/DTLS/UDP) or dedicated RTP streams with proprietary payload format to transmit metadata from a client to a sender.
  • the client may have adjusted into a new pose. To ensure a better quality of experience, the client could adapt the received media to the current factual pose. Indicating what metadata information was used to generate the media, is useful for said adaptation. Assuming the generated media is delivered via RTP stream, an RTP header may be used to contain the metadata information back to the client, e.g. the pose information for which a give video frame was generated.
  • RTP real-time protocol
  • Delivery of a remote rendered content is typically done over real-time protocol (RTP) streams.
  • RTP real-time protocol
  • RTP header extensions become redundant for the receiver because multiple RTP streams may be associated with the same header extension data, e.g. the same pose was used for generating multiple streams.
  • the audio RTP HE is to be superseded by the video RTP HE pose information. This may occur if the audio stream with the sender generated rendering information is retransmitted by an SFU and/or used by a player for audio visual rendering (instead of audio only rendering) where the audio is intended to be consumed according to the RTP HE pose information of the visual content. Currently, such operation will require deep inspection and pose RTP HE removal from the audio stream.
  • a feature to a session negotiation process is implemented, e.g. using a session description protocol (e.g., SDP), for offer-answer negotiation.
  • This feature may comprise a signaling mechanism during session negotiation that informs a client that the RTP header extension information provided by one media stream may be utilized for processing another media stream.
  • SDP SDP signaling for establishing a streaming session between a sender of a content and a receiver for the content which comprises multiple media and metadata.
  • SDP is used as an example and is not limiting but similar principles may be implemented with other kinds of streaming sessions and protocols.
  • the signaling described by this specification could use any other signaling or transport or application end-to-end protocol.
  • An example of such session is presented in Fig. 1.
  • a client 101 negotiates with a server 102 over the negotiation channel 201, e.g. using an SDP description.
  • the client 101 informs the server 102 that the client 101 will send metadata information 301 to the server 102; the server 102 informs the client 101 that server 102 generates at least two media streams (401, 402) based on the received metadata information 301 and sends those media streams to the client 101. Additionally, the server
  • the server 102 informs the client 101 that the server 102 includes in one of the media streams, for example in the first media stream 401, all or a portion of the metadata information 301 based on which the first media stream 401 and the second media stream 402 were generated.
  • a new parameter also called as an indication parameter or reuse information in the following, can be provided to the client 101.
  • the line including the parameter is written with bold letters to improve discemibility of the parameter definition.
  • the indication parameter is included in extension attributes in the RTP header extension of a media, i.e. ⁇ extensionattributes> of extmap attribute of the media.
  • ⁇ extensionattributes> of extmap attribute of the media.
  • a extmap : ⁇ value> [ " / " ⁇ direction>] ⁇ URI> ⁇ extensionattributes>
  • a first example of the indication parameter could be as follows: media : ⁇ mid-li st>
  • the parameter indicates, using an identifier of media mid, to which other media streams the given header extension is applicable.
  • ⁇ mid-list> is a string containing media id values (mid) separated by a semicolon.
  • the parameter hdrext-reuse indicates, using mid of media, to which other media the header extension with id indicated by ⁇ value> is applicable to.
  • ⁇ mid-list> is a string containing media id values (mid) separated by a semi colon.
  • a third example further extends the solutions from the first and second example by providing additional information that can be applied to the header extension when it is reused by other media.
  • additional information can be:
  • interpolation - interpolation method when pose in one media (e.g. video) does not match all media samples in the other media (e.g. audio).
  • the interpolation can be provided as a lookup table of a value to a corresponding interpolation method.
  • An example use of the interpolation method can be the following, where the interpolation value may be one of the following set ⁇ 1, 2, 3 ⁇ :
  • value 1 - client uses the last x values to predict the pose for the current time of the media that is re-using the RTP HE (The prediction may require first synchronizing the two media);
  • offset - indicating default pose translation between one media (e.g. audio) that provides the RTP HE and other media (e.g. video) that reuses the RTP HE.
  • media e.g. audio
  • other media e.g. video
  • the offset can be a delta value that when added to the pose that applies to one stream provides the pose that applies to the other.
  • the offset can be indicated by x,y,z offset values corresponding to each axis of a coordinate system.
  • the translation values, x, y, z may be defined as a function of time drift between the last received pose value (e.g., in an audio stream) and the current media sample of the dependent stream (e.g., the video stream). This method may be more suited to media streams that are synchronized at the receiver.
  • An example of the indication parameter on top attributes presented in first example can be as follows: media : ⁇ mid-list> interpolation : ⁇ value> pof f set : ⁇ value> [0108]
  • extension URI/ID implicitly informs that the extension is applicable to all media that are in the same bundle as media that contains the RTP HE.
  • the BUNDLE grouping has additional parameters that can inform about the interpolation as described above in the third example.
  • ⁇ value> provides the header extension id; ⁇ mid-list> informs to which other media the header extension with id indicated by ⁇ value> is applicable to, using mid of media; ⁇ interpolation> provides the interpolation method information.
  • a hdrext-of f set : ⁇ value> ⁇ mid-list> ⁇ of fset>
  • ⁇ value> provides the header extension id
  • ⁇ mid-list> informs to which other media the header extension with id indicated by ⁇ value> is applicable to, using mid of media
  • ⁇ offset > provides the offset that should be applied to the header information when applied to the media listed in ⁇ mid-list>
  • the indication parameter for header extension reuse has the additional semantic to indicate that synchronization between the two media streams is required.
  • RTCP mechanisms may be used to first synchronize the timing of the two or more streams to adequately map the information received in the RTP HE of one, e.g., pose data for remote rendering, to the content received in the other stream(s).
  • Synchronization may be specifically required for timed metadata such as pose information.
  • the synchronization requirement can be signalled using the indication parameters described above.
  • the parameter synch can be similarly used for other examples provided above.
  • this scenario may relate to conversational applications in which the server 102 operates as a selecting forwarding unit (SFU) or a multipoint control unit (MCU).
  • SFU selecting forwarding unit
  • MCU multipoint control unit
  • the SFU may receive media from each party in a conference call, decides which streams should be forwarded to other parties, and then forwards them to one or more other party of the conference call.
  • the MCU allows for multi-party communications, wherein the MCU may receive media from each party and forward the media to other parties of the multi-party communications.
  • Fig. 2 illustrates this scenario in a simplified manner, showing only two clients 101, 103 and the server 102.
  • the client 101 negotiates with the server 102 over the negotiation channel 201 e.g. using SDP description.
  • the negotiation between client 101 and server 102 contains all or part of the following information: the client 101 informs the server 102 that the client 101 sends metadata information 301 to the server 102; the server 102 receives audio stream from another client 103 with pose information and the server generates media streams (401, 402) based on the received metadata information 301 and the received audio stream with legacy pose information 403 and sends those media streams back to the client 101, where the audio stream is forwarded by the server 102 without further metadata modification; the server 102 includes in one of the media streams (401) all or a portion of the metadata information 301 based on which the media streams were generated; the client 101 knows that the metadata information 301 included in media streams (401) is applicable to other media stream (402).
  • An MTSI receiver Multimedia Telephony Services over IP Multimedia Subsystem, IMS
  • IMS IP Multimedia Subsystem
  • That ROI MTSI sender signals the sent ROI by inserting the HE "um:3gpp:roi-sent" only in one of the RTP streams, and signals the HE to be re-used for the other component streams.
  • the header extension associated with the first media stream ml is also associated with the second media stream m2 (attribute component) and the third media stream m3 (atlas component).
  • an RTP header extension is carried with an RTP stream, referred to here as the independent stream, and utilized by one or more additional RTP streams, referred to here as the dependent streams.
  • the RTP header extension is defined such that the independent RTP stream carries the full pose information and the dependent RTP streams carry only the offset in the header extensions.
  • the first mid in a group :bundle is the independent stream, and the others that carry the extension are the dependent streams.
  • the BUNDLE extension RFC8843 is used to signal RTP sessions containing multiple media types.
  • SDP can be utilized in a scenario where two entities (i.e., server and client) negotiate at a common understanding and setup of a multimedia session between them.
  • one entity e.g., server
  • This offer/answer model is described in RFC3264 and can be used by other protocols for example Session Initiation Protocol (SIP) RFC 3264.
  • SIP Session Initiation Protocol
  • V3C bitstream is a composition of number of video sub-bitstreams and atlas subbitstreams that can be identified by V3C unit headers.
  • Video sub-bitstream can be coded by well-known video coding standards such as AVC, HEVC, VVC which are NAE unit based and have well defined RTP payload format (RFC 6184, RFC 7798, and RFC 9328, respectively).
  • V3C Parameter Set and V3C Unit headers As each of the sub-bitstream of V3C has its own RTP format type and can be separately represented by the SDP, the problem arises on how to indicate the relation between separate RTP sessions and how to pass to a receiver the information that is provided by V3C Parameter Set and V3C Unit headers.
  • Fig. 3 illustrates a possible architecture on how a V3C bitstream can be delivered over network using RTP protocols.
  • a V3C bitstream is provided 403 to the server 102 by the second client 103.
  • the server 102 is configured to demultiplex the V3C bitstream into number of V3C sub-bitstreams.
  • Each V3C sub-bitstream may be encapsulated to an appropriate RTP payload format and sent over a dedicated RTP session to the client 101.
  • the server 102 creates a session specific file format, such as a SDP file, that describes each RTP session as well as provides the above described information that allows the client 101 to reuse some of the parameters of one media with one or more other media.
  • a V3C decoder/renderer 104 Using the information provided by the SDP and over RTP session the client 101 is able to reconstruct V3C bitstream and provide 404 it to a V3C decoder/renderer 104.
  • SDP can be provided in a declarative manner where a client does not have any decision capability, or in case a server has capability to re-encode V3C sub-bitstreams, or V3C sub-bitstreams are provided in number of alternatives, the SDP may be used in offer/ answer mode to allow a client to choose the most appropriate codecs.
  • the offer example shows a session description with one V3C object that is composed of four media descriptions that correspond to four V3C components: occupancy, geometry, attribute, and atlas.
  • the video components of V3C are proposed to a client in three different coding alternatives H.264, H.265, and H.266.
  • the method generally comprises forming 501 a session description containing information about at least a first media stream and a second media stream; receiving 502 at least the first media stream and the second media stream of a coded volumetric video; associating 503 the first media stream a real time transport packet header extension; optionally leaving 504 the second media stream without a real time transport packet header extension; and providing 505 information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
  • Each of the steps can be implemented by a respective module of a computer system.
  • the method generally comprises receiving 601 a session description comprising information about at least a first media stream and a second media stream of a coded volumetric video; determining 602 that the first media stream is associated with a real time transport packet header extension; optionally determining 603 that the second media stream is not associated with a real time transport packet header extension; parsing 604 the information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and using 605 the real time transport packet header extension of the first media stream for processing the second media stream.
  • Each of the steps can be implemented by a respective module of a computer system.
  • the apparatus is, for example, the server 102.
  • the apparatus comprises means for receiving 111 at least a first media stream and a second media stream of a coded volumetric video; means for forming 112 a session description containing information about the at least the first media stream and the second media stream; means for associating 113 the first media stream a real time transport packet header extension; means for leaving 114 the second media stream without a real time transport packet header extension; and means for providing 115 information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
  • the means comprises at least one processor, and a memory including a computer program code, wherein the processor may further comprise processor circuitry.
  • the memory and the computer program code are configured to, with the at least one processor, cause the apparatus to perform the method of Fig. 5 according to various embodiments.
  • FIG. 8 An apparatus according to another embodiment is shown in Fig. 8.
  • the apparatus is, for example, the client 101.
  • the apparatus comprises means for receiving 121 a session description containing information about at least a first media stream and a second media stream of a coded volumetric video; means for determining 122 that the first media stream is associated with a real time transport packet header extension; means for determining 123 that the second media stream is not associated with a real time transport packet header extension; means for parsing 124 the information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and means for using 125 the a real time transport packet header extension of the first media stream for processing the second media stream.
  • the means comprises at least one processor, and a memory including a computer program code, wherein the processor may further comprise processor circuitry.
  • the memory and the computer program code are configured to, with the at least one processor, cause the apparatus to perform the method of Fig. 6 according to various embodiments.
  • Fig. 9 shows a block diagram of a video coding system according to an example embodiment as a schematic block diagram of an electronic device 50, which may incorporate a codec.
  • the electronic device may comprise an encoder or a decoder.
  • the electronic device 50 may for example be a mobile terminal or a user equipment of a wireless communication system or a camera device.
  • the electronic device 50 may be also comprised at a local or a remote server or a graphics processing unit of a computer.
  • the device may be also comprised as part of a head-mounted display device.
  • the apparatus 50 may comprise a display 32 in the form of a liquid crystal display.
  • the display may be any suitable display technology suitable to display an image or video.
  • the apparatus 50 may further comprise a keypad 34.
  • any suitable data or user interface mechanism may be employed.
  • the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
  • the apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input.
  • the apparatus 50 may further comprise an audio output device which in embodiments of the invention may be any one of: an earpiece 38, speaker, or an analogue audio or digital audio output connection.
  • the apparatus 50 may also comprise a battery (or in other embodiments of the invention the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator).
  • the apparatus may further comprise a camera 42 capable of recording or capturing images and/or video.
  • the camera 42 may be a multi-lens camera system having at least two camera sensors.
  • the camera is capable of recording or detecting individual frames which are then passed to the codec 54 or the controller for processing.
  • the apparatus may receive the video and/or image data for processing from another device prior to transmission and/or storage.
  • the apparatus 50 may comprise a controller 56 or processor for controlling the apparatus 50.
  • the apparatus or the controller 56 may comprise one or more processors or processor circuitry and be connected to memory 58 which may store data in the form of image, video and/or audio data, and/or may also store instructions for implementation on the controller 56 or to be executed by the processors or the processor circuitry.
  • the controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and decoding of image, video and/or audio data or assisting in coding and decoding carried out by the controller.
  • the apparatus 50 may further comprise a card reader 48 and a smart card 46, for example a UICC (Universal Integrated Circuit Card) and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.
  • the apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network.
  • the apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and for receiving radio frequency signals from other apparatus(es).
  • the apparatus may comprise one or more wired interfaces configured to transmit and/or receive data over a wired connection, for example an electrical cable or an optical fiber connection.
  • a device may comprise circuitry and electronics for handling, receiving, and transmitting data, computer program code in a memory, and a processor that, when running the computer program code, causes the device to carry out the features of an embodiment.
  • a network device like a server may comprise circuitry and electronics for handling, receiving, and transmitting data, computer program code in a memory, and a processor that, when running the computer program code, causes the network device to carry out the features of various embodiments.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Data Exchanges In Wide-Area Networks (AREA)

Abstract

The embodiments relate to a method and technical equipment for volumetric video coding. The method comprises forming a session description containing information the at least a first media stream and a second media stream; receiving at least the first media stream and the second media stream; associating the first media stream a real time transport packet header extension; and providing reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.

Description

A METHOD, AN APPARATUS AND A COMPUTER PROGRAM PRODUCT FOR
VIDEO CODING
Technical Field
[0001] The present solution generally relates to coding of volumetric video.
Background
[0002] Volumetric video data represents a three-dimensional (3D) scene or object, and can be used as input for AR (Augmented Reality), VR (Virtual Reality), and MR (Mixed Reality) applications. Such data describes geometry (Shape, size, position in 3D space) and respective attributes (e.g., color, opacity, reflectance, . ..), and any possible temporal transformations of the geometry and attributes at given time instances (like frames in 2D video). Volumetric video can be generated from 3D models, also referred to as volumetric visual objects, i.e., CGI (Computer Generated Imagery), or captured from real-world scenes using a variety of capture solutions, e.g., multi-camera, laser scan, combination of video and dedicated depth sensors, and more. Also, a combination of CGI and real-world data is possible.
[0003] Considering that visual volumetric video-based coding (V3C) compresses volumetric content using video codecs, traditional video coding systems can be leveraged. As such, video components containing occupancy, geometry or attribute information may be encapsulated into Real-time Transport Protocol (RTP) video streams. There is also ongoing work to define RTP payload format for atlas data, which would enable streaming of volumetric content over multiple RTP streams.
[0004] For audio-visual scenes with complex 3D geometry or visual and audio properties the rendering of both audio and visual experience on a light user end-point devices may become impractical due to limited hardware capabilities, heat dissipation and battery constraints. Remote rendering can be used to partially solve the problem, by shifting part or all the processing on the network.
Summary
[0005] The scope of protection sought for various embodiments of the invention is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the invention. [0006] Various aspects include a method, an apparatus and a computer readable medium comprising a computer program stored therein, which are characterized by what is stated in the independent claims. Various embodiments are disclosed in the dependent claims.
[0007] In accordance with some embodiments the problem presented above may be solved by introducing a new feature to a session negotiation process, e.g. in using the session description protocol (e.g., SDP) for offer-answer negotiation. The new feature would consist of a signaling mechanism during session negotiation that informs a client that a RTP header extension information provided by one media stream may be utilized for processing another media stream.
[0008] In one embodiment a new signalling is provided by additional attributes in the RTP header extension uniform resource identifier (URI) / identifier (ID) to media mapping.
[0009] In another embodiment such signalling would be implicitly provided by the extension definition, i.e. URI or ID, that a given RTP header extension is applicable to all media present in the SDP group :bundle.
[0010] In another embodiment such signaling could be provided in the RTP header extension itself utilizing the media id MID.
[0011] In another embodiment, the RTP header extension is attached to multiple streams at different frames. This signalling would indicate that the header extension is flexibly present in different streams.
[0012] In accordance with an embodiment a method is defined which produces the following
- a session description document containing information about two or more media streams
- wherein the first media stream contains RTP header extensions
- wherein the second or more media streams do not contain any RTP header extension
- the session description document contains information indicating that the RTP header extension of the first media stream can be utilized for the processing of the second media stream. [0013] In accordance with an embodiment a method is defined which parses the following
- a session description document containing information about two or more media streams
- wherein the first media stream contains RTP header extensions
- wherein the second media stream does not contain any RTP header extension
- parses the information indicating that the RTP header extension of the first media stream can be utilized for the processing of the second media stream
- uses the RTP header extension of the first media stream for processing the second media stream.
[0014] According to a first aspect, there is provided an apparatus comprising means for forming a session description containing information about at least a first media stream and a second media stream; means for receiving at least the first media stream and the second media stream; means for associating the first media stream a real time transport packet header extension; and means for providing reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
[0015] According to a second aspect, there is provided a method, comprising forming a session description containing information about at least a first media stream and a second media stream; receiving at least the first media stream and the second media stream; associating the first media stream a real time transport packet header extension; and providing reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
[0016] According to a third aspect, there is provided an apparatus comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following: forming a session description containing information about at least a first media stream and a second media stream; receiving at least the first media stream and the second media stream; associating the first media stream a real time transport packet header extension; and providing reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream. [0017] According to a fourth aspect, there is provided computer program product comprising computer program code configured to, when executed on at least one processor, cause an apparatus or a system to form a session description containing information about at least a first media stream and a second media stream; receive at least the first media stream and the second media stream; associate the first media stream a real time transport packet header extension; and provide reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
[0018] According to an embodiment, the second media stream is left without a real time transport packet header extension.
[0019] According to an embodiment, the reuse information is provided in the real time transport packet header extension.
[0020] According to an embodiment, identification of the second media stream is provided in the reuse information.
[0021] According to an embodiment, the real time transport packet header extension is attached to multiple streams at different frames, wherein the reuse information indicates that the real time transport packet header extension is flexibly present in different streams. [0022] According to an embodiment, the reuse information is included in extension attributes in the real time transport packet header extension.
[0023] According to an embodiment, the reuse information is expresses as a parameter a=hdrext-reuse : <value> <mid-li st>, in which <value> indicates which other media the header extension is applicable to, and <mid- li st> contains identifiers of the other media.
[0024] According to an embodiment, the first media stream and the second media stream are streams of different components of a volumetric video.
[0025] According to an embodiment, the computer program product is embodied on a non-transitory computer readable medium.
[0026] According to a fifth aspect, there is provided an apparatus comprising means for receiving a session description containing information about at least a first media stream and a second media stream; means for determining that the first media stream is associated with a real time transport packet header extension; means for parsing a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and means for using the real time transport packet header extension of the first media stream for processing the second media stream.
[0027] According to a sixth aspect, there is provided a method, comprising receiving a session description containing information about at least a first media stream and a second media stream; determining that the first media stream is associated with a real time transport packet header extension; parsing a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and using the real time transport packet header extension of the first media stream for processing the second media stream.
[0028] According to a seventh aspect, there is provided an apparatus comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following: receiving a session description containing information about at least a first media stream and a second media stream; determining that the first media stream is associated with a real time transport packet header extension; parsing a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and using the real time transport packet header extension of the first media stream for processing the second media stream. [0029] According to an eighth aspect, there is provided computer program product comprising computer program code configured to, when executed on at least one processor, cause an apparatus or a system to receive a session description containing information about at least a first media stream and a second media stream; determine that the first media stream is associated with a real time transport packet header extension; parse a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and use the real time transport packet header extension of the first media stream for processing the second media stream.
[0030] According to an embodiment, the second media stream is left without a real time transport packet header extension.
[0031] According to an embodiment, the reuse information is parsed from the real time transport packet header extension. [0032] According to an embodiment, identification of the second media stream is parsed from the reuse information.
[0033] According to an embodiment, the real time transport packet header extension is attached to multiple streams at different frames, wherein the reuse information indicates that the real time transport packet header extension is flexibly present in different streams.
[0034] According to an embodiment, the reuse information is received in extension attributes in the real time transport packet header extension.
[0035] Some embodiments may allow to avoid duplication of information in a transmission system. Furthermore, some embodiments enable over-riding legacy header extension information without the need for deep inspection and manipulation of the header extension.
Description of the Drawings
[0036] In the following, various embodiments will be described in more detail with reference to the appended drawings, in which
[0037] Fig. 1 shows an example of a session initiation between a server and a client;
[0038] Fig. 2 illustrates a scenario in which a server receives media from another entity such as another client device;
[0039] Fig. 3 illustrates a possible architecture on how a V3C bitstream can be delivered over network using RTP protocols;
[0040] Fig. 4 shows an example of an extension header;
[0041] Fig. 5 is a flowchart illustrating a method according to an embodiment;
[0042] Fig. 6 is a flowchart illustrating a method according to another embodiment;
[0043] Fig. 7 shows an apparatus according to an embodiment;
[0044] Fig. 8 shows an apparatus according to another embodiment; and
[0045] Fig. 9 shows an apparatus according to an embodiment.
Description of Example Embodiments
[0046] The following description and drawings are illustrative and are not to be construed as unnecessarily limiting. The specific details are provided for a thorough understanding of the disclosure. However, in certain instances, well-known or conventional details are not described in order to avoid obscuring the description. References to one or an embodiment in the present disclosure can be, but not necessarily are, reference to the same embodiment and such references mean at least one of the embodiments.
[0047] Reference in this specification to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment in included in at least one embodiment of the disclosure.
[0048] The present embodiments are in particularly targeted to a solution for reusing a header extension delivered by one media by other media in a bundle.
[0049] There are alternatives to capture and represent a volumetric frame. The format used to capture and represent the volumetric frame depends on the process to be performed on it, and the target application using the volumetric frame. As a first example a volumetric frame can be represented as a point cloud. A point cloud is a set of unstructured points in 3D space, where each point is characterized by its position in a 3D coordinate system (e.g., Euclidean), and some corresponding attributes (e.g., color information provided as RGBA value, or normal vectors). As a second example, a volumetric frame can be represented as images, with or without depth, captured from multiple viewpoints in 3D space. In other words, the volumetric video can be represented by one or more view frames (where a view is a projection of a volumetric scene on to a plane (the camera plane) using a real or virtual camera with known/ computed extrinsic and intrinsic). Each view may be represented by a number of components (e.g., geometry, color, transparency, and occupancy picture), which may be part of the geometry picture or represented separately. As a third example, a volumetric frame can be represented as a mesh. Mesh is a collection of points, called vertices, and connectivity information between vertices, called edges. Vertices along with edges form faces. The combination of vertices, edges and faces can uniquely approximate shapes of objects.
[0050] Depending on the capture, a volumetric frame can provide viewers the ability to navigate a scene with six degrees of freedom, i.e., both translational and rotational movement of their viewing pose (which includes yaw, pitch, and roll). The data to be coded for a volumetric frame can also be significant, as a volumetric frame can contain many numbers of objects, and the positioning and movement of these objects in the scene can result in many dis-occluded regions. Furthermore, the interaction of the light and materials in objects and surfaces in a volumetric frame can generate complex light fields that can produce texture variations for even a slight change of pose. [0051] A sequence of volumetric frames is a volumetric video. Due to large amount of information, storage and transmission of a volumetric video requires compression. A way to compress a volumetric frame can be to project the 3D geometry and related attributes into a collection of 2D images along with additional associated metadata. The projected 2D images can then be coded using 2D video and image coding technologies, for example ISO/IEC 14496-10 (H.264/AVC) and ISO/IEC 23008-2 (H.265/HEVC). The metadata can be coded with technologies specified in specification such as ISO/IEC 23090-5. The coded images and the associated metadata can be stored or transmitted to a client that can decode and render the 3D volumetric frame.
[0052] A Real time Transport Protocol (RTP) is intended for an end-to-end, real-time transport or streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP allows data transport to multiple destinations through IP multicast or to a specific destination through IP unicast. The majority of the RTP implementations are built on the User Datagram Protocol (UDP). Other transport protocols may also be utilized. RTP is used in together with other protocols such as H.323 and Real Time Streaming Protocol RTSP.
[0053] The RTP specification describes two protocols: RTP and RTCP. RTP is used for the transport of multimedia data, and the RTCP (RTP Control Protocol) is used to periodically send control information and QoS parameters.
[0054] RTP sessions may be initiated between client and server using a signalling protocol, such as H.323, the Session Initiation Protocol (SIP), or RTSP. These protocols may use the Session Description Protocol (RFC 8866) to specify the parameters for the sessions.
[0055] RTP is designed to carry a multitude of multimedia formats, which permits the development of new formats without revising the RTP standard. To this end, the information required by a specific application of the protocol is not included in the generic RTP header. For a class of applications (e.g., audio, video), an RTP profile may be defined. For a media format (e.g., a specific video coding format), an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may require a profile and payload format specifications. [0056] The profile defines the codecs used to encode the payload data and their mapping to payload format codecs in the protocol field Payload Type (PT) of the RTP header.
[0057] For example, RTP profile for audio and video conferences with minimal control is defined in RFC 3551. The profile defines a set of static payload type assignments, and a dynamic mechanism for mapping between a payload format, and a PT value using Session Description Protocol (SDP). The latter mechanism is used for newer video codec such as RTP payload format for H.264 Video defined in RFC 6184 or RTP Payload Format for High Efficiency Video Coding (HEVC) defined in RFC 7798.
[0058] An RTP session is established for each multimedia stream. Audio and video streams may use separate RTP sessions, enabling a receiver to selectively receive components of a particular stream. The RTP specification recommends even port number for RTP, and the use of the next odd port number of the associated RTCP session. A single port can be used for RTP and RTCP in applications that multiplex the protocols.
[0059] Each RTP stream consists of RTP packets, which in turn consist of RTP header and payload pairs.
[0060] RTP packets are created at the application layer and handed to the transport layer for delivery. Each unit of RTP media data created by an application begins with the RTP packet header.
[0061] The RTP header has a minimum size of 12 bytes. After the header, optional header extensions may be present. This is followed by the RTP payload, the format of which is determined by the particular class of application. The fields in the header are as follows:
• Version: (2 bits) Indicates the version of the protocol.
• P (Padding): (1 bit) Used to indicate if there are extra padding bytes at the end of the RTP packet.
• X (Extension): (1 bit) Indicates the presence of an extension header between the header and payload data. The extension header is application or profile specific.
• CC (CSRC count): (4 bits) Contains the number of CSRC identifiers that follow the SSRC • M (Marker) : (1 bit) Signalling used at the application level in a profile-specific manner. If it is set, it means that the current data has some special relevance for the application.
• PT (Payload type): (7 bits) Indicates the format of the payload and thus determines its interpretation by the application.
• Sequence number: (16 bits) The sequence number is incremented for each RTP data packet sent and is to be used by the receiver to detect packet loss and to accommodate out-of-order delivery.
• Timestamp: (32 bits) Used by the receiver to play back the received samples at appropriate time and interval. When several media streams are present, the timestamps may be independent in each stream. The granularity of the timing is application specific. For example, video stream may use a 90 kHz clock. The clock granularity is one of the details that is specified in the RTP profile for an application.
• SSRC: (32 bits) Synchronization source identifier uniquely identifies the source of the stream. The synchronization sources within the same RTP session will be unique.
• CSRC: (32 bits each) Contributing source IDs enumerate contributing sources to a stream which has been generated from multiple sources.
• Header extension: (optional, presence indicated by Extension field) The first 32-bit word contains a profile-specific identifier (16 bits) and a length specifier (16 bits) that indicates the length of the extension in 32-bit units, excluding the 32 bits of the extension header. The extension header data is shown in Fig. 4.
[0062] The RTP specification RFC3550 provides a capability to extend the RTP header. Section 5.3.1 of RFC3550 defines the header extension format and rules for its use. Header extensions may carry metadata in addition to the usual RTP header information, provided the RTP layer can function if that metadata is missing. The RTP specification RFC3550 states that RTP "is designed so that the header extension may be ignored by other interoperating implementations that have not been extended." The intent of this restriction is that RTP header extensions (RTP HE) must not be used to extend RTP itself in a manner that is backward incompatible with non-extended implementations. [0063] Additionally, the general mechanism for HE specification, RFC8285, provides the option to use a small number of extensions in each RTP packet, where the domain of possible extensions is large and registration is decentralized.
[0064] The actual extensions in use in a RTP session are signalled in the setup information for that session.
[0065] In this disclosure, the Session Description Protocol (SDP) is used as an example of a session specific file format. SDP is a format for describing multimedia communication sessions for the purposes of announcement and invitation. Its predominant use is in support of conversational and streaming media applications. SDP does not deliver any media streams itself, but is used between endpoints for negotiation of network metrics, media types, bandwidth requirements, and other associated properties. The set of properties and parameters is called a session profile. SDP is extensible for the support of new media types and formats.
[0066] The Session Description Protocol describes a session as a group of fields in a text-based format, one field per line. The form of each field is as follows: <character>=<value><CR><LF> where <character> is a single case-sensitive character and <value> is structured text in a format that depends on the character. Values may be UTF-8 encoded. Whitespace is not allowed immediately to either side of the equal sign.
[0067] Session descriptions consist of three sections: session, timing, and media descriptions. Each description may contain multiple timing and media descriptions. Names are only unique within the associated syntactic construct.
[0068] Fields appear in the order, shown below; optional fields are marked with an asterisk: v= (protocol version number, currently only 0 ) o= ( originator and ses sion identi fier : username , id, version number, network address ) s= ( ses sion name : mandatory with at least one UTF- 8-encoded character ) i=* ( session title or short information ) u=* (URI of description ) e=* ( zero or more email address with optional name of contacts ) p=* ( zero or more phone number with optional name of contacts ) c=* (connection inf ormation— not required if included in all media) b=* (zero or more bandwidth information lines)
One or more time descriptions ("t=", "r=" and "z=" lines; see below) z=* (time zone adjustments) k=* (obsolete) a=* (zero or more session attribute lines)
Zero or more Media descriptions (each one starting by an "m=" line; see below)
[0069] Time description (mandatory): t= (time the session is active) r=* (zero or more repeat times) z=* (optional time zone offset line)
[0070] Media description (optional): m= (media name and transport address) i=* (media title or information field) c=* (connection information — optional if included at session level) b=* (zero or more bandwidth information lines) k=* (obsolete) a=* (zero or more media attribute lines — overriding the Session attribute lines)
[0071] Below is a sample session description from RFC 8866. This session is originated by the user "jdoe", at IPv4 address 198.51.100.1. Its name is "Call to John Smith" and the session information ("SDP Offer #1") is included along with a link for additional information as well as a phone number and an email address to contact the responsible party, Jane Doe. This session v=0 o=jdoe 3724394400 3724394405 IN IP4 198.51.100.1 s=Call to John Smith i=SDP Offer #1 u=http : / / www . doe . example . com/home . html e=Jane Doe < ane@ doe . example . com> p=+l 617 555-6011 c=IN IP4 198.51.100.1 t=0 0 m=audio 49170 RTP/AVP 0 m=audio 49180 RTP/AVP 0 m=video 51372 RTP/AVP 99 c=IN IP6 2001 : db8 : : 2 a=rtpmap : 99 h263-1998 / 90000
[0072] SDP uses attributes to extend the core protocol. Attributes can appear within the Session or Media sections and are scoped accordingly as session-level or media-level. New attributes can be added to the standard through registration with IANA. A media description may contain any number of “a=” lines (attribute-fields) that are media description specific. Session-level attributes convey additional information that applies to the session as a whole rather than to individual media descriptions.
[0073] Attributes are either properties or values: a=< at tribute- name > a=<attribute-name> : <attribute-value>
[0074] Examples of attributes defined in RFC8866 are “rtpmap” and “fimtp”. “rtpmap” attribute maps from an RTP payload type number (as used in an "m=" line) to an encoding name denoting the payload format to be used. It also provides information on the clock rate and encoding parameters. Up to one "a=rtpmap:" attribute can be defined for each media format specified. This can be the following: m=audio 49230 RTP/AVP 96 97 98 a=rtpmap : 96 L8 / 8000 a=rtpmap : 97 L16/ 8000 a=rtpmap : 98 L16/ 11025/ 2
[0075] In the example above, the media types are “audio/L8” and “audio/L16”. Parameters added to an "a=rtpmap:" attribute may only be those required for a session directory to make the choice of appropriate media to participate in a session. Codecspecific parameters may be added in other attributes, for example, "fimtp".
[0076] "fimtp" attribute allows parameters that are specific to a particular format to be conveyed in a way that SDP does not have to understand them. The format can be one of the formats specified for the media. Format-specific parameters, semicolon separated, may be any set of parameters required to be conveyed by SDP and given unchanged to the media tool that will use this format. At most one instance of this attribute is allowed for each format. An example is: a=fmtp : 96 profile-level-id=42e016 ; max-mbps=108000 ; max-fs=3600
[0077] For example RFC7798 defines the following sprop-vps, sprop-sps, sprop-pps, profile-space, profile-id, tier-flag, level-id, interop-constraints, profile-compatibility- indicator, sprop-sub-layer-id, recv-sub-layer-id, max-recv-level-id, tx-mode, max-lsr, max- Ips, max-cpb, max-dpb, max-br, max-tr, max-tc, max-fps, sprop-max-don-diff, sprop- depack-buf-nalus, sprop-depack-buf-bytes, depack-buf-cap, sprop-segmentation-id, sprop- spatial-segmentation-idc, dec-parallel-cap, and include-dph.
[0078] ‘ ‘group” and “mid” attributes defined in RFC 5888 allows to group "m" lines in SDP for different purposes. An example can be for lip synchronization or for receiving a media flow consisting of several media streams on different transport addresses.
[0079] An example would be in a given session description, each "m" line is identified by a token, which is carried in a "mid" attribute below the "m" line. The session description carries session-level "group" attributes that group different "m" lines (identified by their tokens) using different group semantics. The semantics of a group describe the purpose for which the "m" lines are grouped. In the example below, the "group" line indicates that the "m" lines identified by tokens 1 and 2 (the audio and the video "m" lines, respectively) are grouped for the purpose of lip synchronization (LS). v=0 o=Laura 289083124 289083124 IN I P4 one . example . com c=IN IP4 192 . 0 . 2 . 1 t=0 0 a=group : LS 1 2 m=audio 30000 RTP/AVP 0 a=mid : 1 m=video 30002 RTP/AVP 31 a=mid : 2
[0080] RFC5888 defines two semantics for group Lip Synchronization (LS), as used in the example above, and Flow Identification (FID). RFC5583 defines another grouping type Decoding Dependency (DDP). RFC8843 defines another grouping type BUNDLE, which among other is utilized when multiple types of media are sent in a single RTP session as described in RFC8860.
[0081] Two media-level and one session-level attributes are used in a mechanism for providing alternative SDP lines. One or more SDP lines at media level can be replaced, if desired, by alternatives. The mechanism is backwards compatible in the way that a receiver that does not support the attributes may get the default configuration. The different alternatives can be grouped using different attributes that can be specified hierarchically with a top and a lower level. 3 GPP Release 6 supports grouping based on bit-rate, according to the SDP bandwidth modifiers AS [RFC8866] and TIAS [RFC3890], and language.
[0082] The SDP attributes are:
• The media-level attribute "a=alt:<id>:<SDP-Line>" carries any SDP line and an alternative identifier.
• The media-level attribute "a=alt-default-id:<id>" identifies the default configuration to be used in groupings.
• The session-level attribute "a=alt-group" is used to group different recommended media alternatives. This allows providing aggregated properties for the whole group according to the grouping type. Language and bit-rate are two defined grouping types.
[0083] Each local identifier of RTP HE potentially used in the RTP stream is mapped to an extension identified by a URI using an SDP attribute of the form: a=extmap : <value> [ " / "<direction>] <URI> <extensionattributes> where
• <value> is the local identifier (ID) of this extension and is an integer in the valid range (0 is reserved for padding in both forms, and 15 is reserved in the one-byte header form, as noted above).
• <direction> is one of "sendonly", "recvonly", "sendrecv", or "inactive" (without the quotes) with relation to the device being configured.
• <URI> is a URI, as above
• <extensionattributes> is byte-string [RFC 8866] defined by the mapped extension [0084] An example usage of the attribute for a give media is presented below: m=video a=sendrecv a=extmap : l URI -toff set a=extmap : 3 URI -f rametype m=audio a=sendrecv a=extmap : l URI -toff set
[0085] The SDP attribute extmap is a SPECIAL category SDP attribute as defined in RFC 8859. For the attributes in the SPECIAE category, the text in the specification defining the attribute must be consulted for further handling when multiplexed. Therefore, multiplexing considerations for RTP header extensions need to be defined by the specification of every RTP header extension.
[0086] Remote Rendering offloads rendering process of 3D content to a cloud and streams it to devices in real time. Due to that an end user may be able to use interactive, high-quality 3D models with every detail intact and no compromise in quality on a devices with limited capabilities.
[0087] In a remote rendering scenario a client may send continuous metadata information to a server based on which appropriate media is provided to the client. E.g. a client may send its pose in its local virtual environment, based on the pose the server renders the 3D scene into 2D views, and subsequently encodes the 2D views into a video which can be streamed to the client. For example, gaming platforms utilize typically WebRTC data channels (SCTP/DTLS/UDP) or dedicated RTP streams with proprietary payload format to transmit metadata from a client to a sender.
[0088] By the time the client receives the media that was generated at the sender, the client may have adjusted into a new pose. To ensure a better quality of experience, the client could adapt the received media to the current factual pose. Indicating what metadata information was used to generate the media, is useful for said adaptation. Assuming the generated media is delivered via RTP stream, an RTP header may be used to contain the metadata information back to the client, e.g. the pose information for which a give video frame was generated.
[0089] Delivery of a remote rendered content is typically done over real-time protocol (RTP) streams. To allow the receiver to perform adjustment on the remote rendered content, e.g. correct/warp the received view based on the information of its current factual pose and the pose that was used to render the view, it is useful to attach the information that was used to render the frames to the delivered frames. This kind of attachment can be conveniently done using RTP header extensions. [0090] When both audio and video are delivered together, or when the delivery of either audio or video is done over multiple real-time streams (e.g. depth+texture), some of the RTP header extensions become redundant for the receiver because multiple RTP streams may be associated with the same header extension data, e.g. the same pose was used for generating multiple streams.
[0091] Furthermore, in some scenarios, the audio RTP HE is to be superseded by the video RTP HE pose information. This may occur if the audio stream with the sender generated rendering information is retransmitted by an SFU and/or used by a player for audio visual rendering (instead of audio only rendering) where the audio is intended to be consumed according to the RTP HE pose information of the visual content. Currently, such operation will require deep inspection and pose RTP HE removal from the audio stream. [0092] In the following, some embodiments will be described in which a feature to a session negotiation process is implemented, e.g. using a session description protocol (e.g., SDP), for offer-answer negotiation. This feature may comprise a signaling mechanism during session negotiation that informs a client that the RTP header extension information provided by one media stream may be utilized for processing another media stream.
[0093] The following disclosure of example embodiments rely on SDP signaling for establishing a streaming session between a sender of a content and a receiver for the content which comprises multiple media and metadata. However, SDP is used as an example and is not limiting but similar principles may be implemented with other kinds of streaming sessions and protocols. For this purpose, the signaling described by this specification could use any other signaling or transport or application end-to-end protocol. An example of such session is presented in Fig. 1.
[0094] In the example of Fig. 1, a client 101 negotiates with a server 102 over the negotiation channel 201, e.g. using an SDP description. The negotiation between the client
101 and the server 102 contains all or part of the following information, in accordance with an embodiment: the client 101 informs the server 102 that the client 101 will send metadata information 301 to the server 102; the server 102 informs the client 101 that server 102 generates at least two media streams (401, 402) based on the received metadata information 301 and sends those media streams to the client 101. Additionally, the server
102 informs the client 101 that the server 102 includes in one of the media streams, for example in the first media stream 401, all or a portion of the metadata information 301 based on which the first media stream 401 and the second media stream 402 were generated.
[0095] In the following, some examples will be provided how a new parameter, also called as an indication parameter or reuse information in the following, can be provided to the client 101. In the examples the line including the parameter is written with bold letters to improve discemibility of the parameter definition.
[0096] In accordance with a first example, the indication parameter is included in extension attributes in the RTP header extension of a media, i.e. <extensionattributes> of extmap attribute of the media. This can be expressed as follows: a=extmap : <value> [ " / "<direction>] <URI> <extensionattributes>
[0097] A first example of the indication parameter could be as follows: media : <mid-li st>
[0098] The parameter indicates, using an identifier of media mid, to which other media streams the given header extension is applicable. <mid-list> is a string containing media id values (mid) separated by a semicolon.
[0099] An example of SDP utilizing this indication parameter is presented below. In the example, the RTP HE with URI um:ietf:params:rtp-hdrext:avt-example-metadata provided in media with mid ml is also applicable to an audio media with mid m2. v=0 o=alice 2890844526 2890844526 IN IP4 host . atlanta . example . com s=AR Ses sion c=IN IP4 host . atlanta . example . com t=0 0 m=application 1001 UDP/DTLS/ SCTP webrtc-datachannel a=sendonly m=video 27568 RTP/AVP 96 a=mid : ml a=recvonly a=rtpmap : 96 H264 / 90000 a=extmap : l urn : ietf : params : rtp-hdrext : avt-example-metadata media : m2 m=audio 27572 RTP/AVP 97 a=mid : m2 a=recvonly a=rtpmap : 0 PCMU/ 8000 m=audio 27574 RTP/AVP 97 a=mid : m3 a=recvonly a=rtpmap:O PCMU/8000 [0100] The parameter media : m2 indicates, using mid of media, to which other media the given header extension is applicable. This parameter is included at the end of the extmap-definition above: a=extmap: 1 um:ietf:params:rtp-hdrext:avt-example-metadata media:m2.
[0101] A second example of the indication parameter could be as follows: a=hdrext-reuse : <value> <mid-list>
[0102] The parameter hdrext-reuse indicates, using mid of media, to which other media the header extension with id indicated by <value> is applicable to. <mid-list> is a string containing media id values (mid) separated by a semi colon.
[0103] An example of SDP utilizing this parameter is presented below. In the example, the RTP HE with URI um:ietf:params:rtp-hdrext:avt-example-metadata provided in media with mid ml is also applicable to an audio media with mid m3. v=0 o=alice 2890844526 2890844526 IN IP4 host.atlanta.example.com sAR Session c=IN IP4 host.atlanta.example.com t=0 0 m=application 1001 UDP/DTLS/SCTP webrtc-datachannel a=sendonly m=video 23458 RTP/AVP 96 a=mid : ml a=recvonly a=rtpmap:96 H264/90000 a=extmap :1 urn:ietf: params : rtp-hdrext : avt- example -metadata a=hdrext-reuse : 1 m3 m=audio 23468 RTP/AVP 97 a=mid : m2 a=recvonly a=rtpmap:97 PCMU/8000 m=audio 23478 RTP/AVP 97 a=mid : m3 a=recvonly a=rtpmap : 97 PCMU/ 8000
[0104] A third example further extends the solutions from the first and second example by providing additional information that can be applied to the header extension when it is reused by other media. Such additional information can be:
• interpolation - interpolation method when pose in one media (e.g. video) does not match all media samples in the other media (e.g. audio). The interpolation can be provided as a lookup table of a value to a corresponding interpolation method. An example use of the interpolation method can be the following, where the interpolation value may be one of the following set {1, 2, 3} :
• value 1 - client uses the last x values to predict the pose for the current time of the media that is re-using the RTP HE (The prediction may require first synchronizing the two media);
• value 2 - client uses the pose of the last received media sample;
• value 3 - client uses the pose to be a constant if the value is missing
• offset - indicating default pose translation between one media (e.g. audio) that provides the RTP HE and other media (e.g. video) that reuses the RTP HE. Another example is when stereoscopic video is provided and the offset provides the pose difference between the two views. Thus, the offset can be a delta value that when added to the pose that applies to one stream provides the pose that applies to the other.
[0105] The offset can be indicated by x,y,z offset values corresponding to each axis of a coordinate system.
[0106] The translation values, x, y, z, may be defined as a function of time drift between the last received pose value (e.g., in an audio stream) and the current media sample of the dependent stream (e.g., the video stream). This method may be more suited to media streams that are synchronized at the receiver.
[0107] An example of the indication parameter on top attributes presented in first example can be as follows: media : <mid-list> interpolation : <value> pof f set : <value> [0108] An example of the indication parameter on top of the attributes presented in the second example can be as follows: a=hdrext-reuse:<value> <mid-list> interpolation : <value> pof f set : <value>
[0109] An example of SDP utilizing those additional attributes is presented below: v=0 o=alice 2890844526 2890844526 IN IP4 host.atlanta.example.com s=SDP Session c=IN IP4 host.atlanta.example.com t=0 0 m=application 1001 UDP/DTLS/SCTP webrtc-datachanel a=sendonly m=video 23456 RTP/AVP 96 a=mid : ml a=recvonly a=rtpmap:96 H264/90000 a=extmap :1 urn:ietf: params : rtp-hdrext : avt- example -metadata a=hdrext-reuse : 1 m3 interpolation:! pof fset : 0.1 ;2.3 ; -23.2 m=audio 23468 RTP/AVP 97 a=mid : m2 a=recvonly a=rtpmap:0 PCMU/8000 m=audio 23478 RTP/AVP 97 a=mid : m3 a=recvonly a=rtpmap:0 PCMU/8000
[0110] In accordance with a fourth example, no additional attributes are provided but the extension URI/ID implicitly informs that the extension is applicable to all media that are in the same bundle as media that contains the RTP HE.
[0111] An example of SDP is presented below where the RTP HE from media ml is also applicable to media m2 as it belongs to the same BUNDLE group. v=0 o=alice 2890844526 2890844526 IN IP4 host.atlanta.example.com s= c=IN IP4 host.atlanta.example.com t=o o a=group : BUNDLE ml m2 m=application 1001 UDP/DTLS/SCTP webrtc-datachanel a=sendonly m=video 10006 RTP/AVP 96 a=mid : ml a=recvonly a=rtpmap:96 H264/90000 a=extmap : 1 urn : ietf : params : rtp-hdrext : avt-example -metadata m=audio 10002 RTP/AVP 97 a=mid : m2 a=recvonly a=rtpmap:0 PCMU/8000 m=audio 10004 RTP/AVP 97 a=mid : m3 a=recvonly a=rtpmap:0 PCMU/8000
[0112] In accordance with a fifth example, the BUNDLE grouping has additional parameters that can inform about the interpolation as described above in the third example. v=0 o=alice 2890844526 2890844526 IN IP4 host.atlanta.example.com s= c=IN IP4 host.atlanta.example.com t=0 0 a=group : BUNDLE ml m2 interpolation : 1 m=application 1001 UDP/DTLS/SCTP webrtc-datachanel a=sendonly m=video 10002 RTP/AVP 96 a=mid : ml a=recvonly a=rtpmap:96 H264/90000 a=extmap : 1 urn : ietf : params : rtp-hdrext : avt-example -metadata m=audio 10004 RTP/AVP 97 a=mid : m2 a=recvonly a=rtpmap : 0 PCMU/ 8000 m=audio 10006 RTP/AVP 97 a=mid : m3 a=recvonly a=rtpmap : 0 PCMU/ 8000 [0113] In accordance with a sixth example, the solutions from the third example could be provided as standalone attributes of a given media and mapped directly to media by which the RTP HE is reused. a=hdrext-interpolation : <value> <mid-li st> <interpolation>
[0114] <value> provides the header extension id; <mid-list> informs to which other media the header extension with id indicated by <value> is applicable to, using mid of media; <interpolation> provides the interpolation method information. a=hdrext-of f set : <value> <mid-list> <of fset>
[0115] <value> provides the header extension id; <mid-list> informs to which other media the header extension with id indicated by <value> is applicable to, using mid of media; < offset > provides the offset that should be applied to the header information when applied to the media listed in <mid-list>
[0116] An example of SDP utilizing these new attributes is presented below. In the example, the RTP HE with URI um:ietf:params:rtp-hdrext:avt-example-metadata provided in media with mid ml is also applicable to an audio media with mid m3 with interpolation 1 and is applicable to an audio with mid m2 with offset (0.1, 2.3, -23.2). v=0 o=alice 2890844526 2890844526 IN I P4 host . atlanta . example . com s= c=IN IP4 host . atlanta . example . com t=0 0 m=application 1001 UDP/DTLS/ SCTP webrtc-datachanel a=sendonly m=video 0 RTP/AVP 96 a=mid : ml a=recvonly a=rtpmap : 96 H264 / 90000 a=extmap : 1 urn : ietf : params : rtp-hdrext : avt- example -metadata a=hdrext-reuse : 1 m2 ;m3 a=hdrext- interpolation : 1 m3 1 a=hdrext-off set : 1 m2 0 . 1 ; 2 . 3 ; -23 . 2 m=audio 0 RTP/AVP 97 a=mid : m2 a=recvonly a=rtpmap : 0 PCMU/ 8000 m=audio 0 RTP/AVP 97 a=mid : m3 a=recvonly a=rtpmap : 0 PCMU/ 8000 [0117] Prior to reusing the metadata provided in the RTP HE of a first media for a second media, time synchronization may be required. In some cases, this may require including the attribute ‘ group :LS’ for the relevant mid values in the SDP along with the indication parameter for header extension reuse. In an embodiment, the indication parameter for header extension reuse has the additional semantic to indicate that synchronization between the two media streams is required. In such a case, RTCP mechanisms may be used to first synchronize the timing of the two or more streams to adequately map the information received in the RTP HE of one, e.g., pose data for remote rendering, to the content received in the other stream(s). Synchronization may be specifically required for timed metadata such as pose information. In another embodiment, the synchronization requirement can be signalled using the indication parameters described above. For example, a=hdrext-reuse:<value> <mid-list> <synch> can be used where synch is a flag such that 0 indicates no synchronization is required for reusing this header extension, and 1 indicates that synchronization is required. The parameter synch can be similarly used for other examples provided above.
[0118] In the following, another scenario will be described. In this scenario the server receives media from another entity such as another client device. As an example, this kind of scenario may relate to conversational applications in which the server 102 operates as a selecting forwarding unit (SFU) or a multipoint control unit (MCU). The SFU may receive media from each party in a conference call, decides which streams should be forwarded to other parties, and then forwards them to one or more other party of the conference call. The MCU allows for multi-party communications, wherein the MCU may receive media from each party and forward the media to other parties of the multi-party communications. [0119] Fig. 2 illustrates this scenario in a simplified manner, showing only two clients 101, 103 and the server 102.
[0120] In this scenario the client 101 negotiates with the server 102 over the negotiation channel 201 e.g. using SDP description. The negotiation between client 101 and server 102 contains all or part of the following information: the client 101 informs the server 102 that the client 101 sends metadata information 301 to the server 102; the server 102 receives audio stream from another client 103 with pose information and the server generates media streams (401, 402) based on the received metadata information 301 and the received audio stream with legacy pose information 403 and sends those media streams back to the client 101, where the audio stream is forwarded by the server 102 without further metadata modification; the server 102 includes in one of the media streams (401) all or a portion of the metadata information 301 based on which the media streams were generated; the client 101 knows that the metadata information 301 included in media streams (401) is applicable to other media stream (402).
[0121] In another scenario, different V3C components are carried over different RTP streams. An MTSI receiver (Multimedia Telephony Services over IP Multimedia Subsystem, IMS) requests a certain region of interest (ROI) for spatial random access. That ROI MTSI sender signals the sent ROI by inserting the HE "um:3gpp:roi-sent" only in one of the RTP streams, and signals the HE to be re-used for the other component streams. v=0 o=alice 2890844526 2890844526 IN I P4 host . atlanta . example . com s= c=IN IP4 host . atlanta . example . com t=0 0 a=group : V3C 1 2 3 4 m=video 40000 RTP/AVP 96 a=rtpmap : 96 H264 / 90000 a=fmtp : 96 v3c-unit-header=10000000 / / occupancy a=mid : ml m=video 40002 RTP/AVP 97 a=rtpmap : 97 H264 / 90000 a=extmap : l urn : 3gpp : roi- sent a=hdrext-reuse : 1 m2 ,m3 a=fmtp : 97 v3c-unit-header=18000000 / / geometry a=mid : m2 m=video 40004 RTP/AVP 98 a=rtpmap : 98 H264 / 90000 a=fmtp : 98 v3c-unit-header=20000000 / / attribute a=mid : m3 m=application 40008 RTP/AVP 100 a=rtpmap : 100 v3c/ 90000 a=fmtp : 100 v3c-unit-header=08000000 ; / / atlas a=mid : m4
[0122] In the example above, the header extension associated with the first media stream ml (geometry component) is also associated with the second media stream m2 (attribute component) and the third media stream m3 (atlas component).
[0123] In the embodiments described above, an RTP header extension is carried with an RTP stream, referred to here as the independent stream, and utilized by one or more additional RTP streams, referred to here as the dependent streams. In accordance with an embodiment, the RTP header extension is defined such that the independent RTP stream carries the full pose information and the dependent RTP streams carry only the offset in the header extensions. In this embodiment, the first mid in a group :bundle is the independent stream, and the others that carry the extension are the dependent streams.
[0124] When using SDP signalling, the BUNDLE extension RFC8843 is used to signal RTP sessions containing multiple media types.
[0125] SDP can be utilized in a scenario where two entities (i.e., server and client) negotiate at a common understanding and setup of a multimedia session between them. In such scenario one entity (e.g., server), offers the other a description of the desired session (or possible option of a session) from their perspective, and the other participant answers with the desired session from their perspective. This offer/answer model is described in RFC3264 and can be used by other protocols for example Session Initiation Protocol (SIP) RFC 3264.
[0126] V3C bitstream is a composition of number of video sub-bitstreams and atlas subbitstreams that can be identified by V3C unit headers. Video sub-bitstream can be coded by well-known video coding standards such as AVC, HEVC, VVC which are NAE unit based and have well defined RTP payload format (RFC 6184, RFC 7798, and RFC 9328, respectively).
[0127] As each of the sub-bitstream of V3C has its own RTP format type and can be separately represented by the SDP, the problem arises on how to indicate the relation between separate RTP sessions and how to pass to a receiver the information that is provided by V3C Parameter Set and V3C Unit headers.
[0128] Fig. 3 illustrates a possible architecture on how a V3C bitstream can be delivered over network using RTP protocols.
[0129] In an example shown in Fig. 3, a V3C bitstream is provided 403 to the server 102 by the second client 103. The server 102 is configured to demultiplex the V3C bitstream into number of V3C sub-bitstreams. Each V3C sub-bitstream may be encapsulated to an appropriate RTP payload format and sent over a dedicated RTP session to the client 101. Along the RTP sessions the server 102 creates a session specific file format, such as a SDP file, that describes each RTP session as well as provides the above described information that allows the client 101 to reuse some of the parameters of one media with one or more other media. Using the information provided by the SDP and over RTP session the client 101 is able to reconstruct V3C bitstream and provide 404 it to a V3C decoder/renderer 104.
[0130] SDP can be provided in a declarative manner where a client does not have any decision capability, or in case a server has capability to re-encode V3C sub-bitstreams, or V3C sub-bitstreams are provided in number of alternatives, the SDP may be used in offer/ answer mode to allow a client to choose the most appropriate codecs.
[0131] An example of an offer/answer procedure is provided below. The offer example shows a session description with one V3C object that is composed of four media descriptions that correspond to four V3C components: occupancy, geometry, attribute, and atlas. The video components of V3C are proposed to a client in three different coding alternatives H.264, H.265, and H.266. v=0 o=svcsrv 289083124 289083124 IN I P4 host . example . com s=V3C S IGNALING t=0 0 c=IN I P4 192 . 0 . 2 . 1 / 127 a=group: BUNDLE 1 2 3 4 a=group:V3C 4 3 2 1 v3c-ptl-level-idc=10;v3c-parameter- set=AF6F00939921878 m=video 40000 RTP/AVP 96 97 98 a=rtpmap:96 H264/90000 a=rtpmap:97 H265/90000 a=rtpmap:98 H266/90000 a=v3cmap: 96 v3c-unit-type=2 ; v3c-vps-id=0 ; v3c-atlas-id=0 a=v3cmap: 97 v3c-unit-type=2 ; v3c-vps-id=0 ; v3c-atlas-id=0 a=v3cmap: 98 v3c-unit-type=2 ; v3c-vps-id=0 ; v3c-atlas-id=0 a=mid : ml a=extmap : 1 urn : ietf : params : rtp-hdrext : sdes : mid a=hdrext-reuse : 1 m2 ;m3 m=video 40000 RTP/AVP 96 97 98 a=rtpmap:96 H264/90000 a=rtpmap:97 H265/90000 a=rtpmap:98 H266/90000 a=v3cmap: 96 v3c-unit-type=3 ; v3c-vps-id=0 ; v3c-atlas-id=0 ; a=v3cmap: 97 v3c-unit-type=3 ; v3c-vps-id=0 ; v3c-atlas-id=0 ; a=v3cmap: 98 v3c-unit-type=3 ; v3c-vps-id=0 ; v3c-atlas-id=0 ; a=mid : m2 m=video 40000 RTP/AVP 96 97 98 a=rtpmap:96 H264/90000 a=rtpmap:97 H265/90000 a=rtpmap:98 H266/90000 a=v3cmap: 96 v3c-unit-type=4 ; v3c-vps-id=0 ; v3c-atlas-id=0 a=v3cmap: 97 v3c-unit-type=4 ; v3c-vps-id=0 ; v3c-atlas-id=0 a=v3cmap: 98 v3c-unit-type=4 ; v3c-vps-id=0 ; v3c-atlas-id=0 a=mid : m3 m=video 40000 RTP/AVP 96 a=rtpmap:96 ATLAS/90000 a=v3cmap: 96 v3c-unit-type=l ; v3c-vps-id=0 ; v3c-atlas-id=0 a=mid : m4 a=extmap : 1 urn : ietf : params : rtp-hdrext : sdes : mid [0132] The client provides an SDP answer, where it selects a different video codec for each V3C video component. v=0 o=svcclt 289083124 289083124 IN IP4 host.example.com s=V3C SIGNALING t=0 0 c=IN IP4 192.0.2.2/127 a=group: BUNDLE 1 2 3 4 a=group:V3C 4 3 2 1 v3c-ptl-level-idc=10;v3c-parameter- set=AF6F00939921878 m=video 0 RTP/AVP 96 a=rtpmap:96 H264/90000 a=v3cmap: 96 v3c-unit-type=2 ; v3c-vps-id=0 ; v3c-atlas-id=0 a=mid : ml a=bundle-only a=extmap : 1 urn : ietf : params : rtp-hdrext : sdes : mid m=video 0 RTP/AVP 97 a=rtpmap:97 H265/90000 a=v3cmap: 97 v3c-unit-type=3 ; v3c-vps-id=0 ; v3c-atlas-id=0 ; a=mid : m2 a=bundle-only m=video 0 RTP/AVP 98 a=rtpmap:98 H266/90000 a=v3cmap: 98 v3c-unit-type=4 ; v3c-vps-id=0 ; v3c-atlas-id=0 a=mid : m3 a=bundle-only m=video 40000 RTP/AVP 96 a=rtpmap:96 ATLAS/90000 a=v3cmap: 96 v3c-unit-type=l ; v3c-vps-id=0 ; v3c-atlas-id=0 a=mid : m4 a=extmap : 1 urn : ietf : params : rtp-hdrext : sdes : mid
[0133] In the previous examples, SDP has been used as an example of a session specific file format. [0134] A method according to an embodiment is shown in Fig. 5. The method generally comprises forming 501 a session description containing information about at least a first media stream and a second media stream; receiving 502 at least the first media stream and the second media stream of a coded volumetric video; associating 503 the first media stream a real time transport packet header extension; optionally leaving 504 the second media stream without a real time transport packet header extension; and providing 505 information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
[0135] Each of the steps can be implemented by a respective module of a computer system.
[0136] The method according to another embodiment is shown in Fig. 6. The method generally comprises receiving 601 a session description comprising information about at least a first media stream and a second media stream of a coded volumetric video; determining 602 that the first media stream is associated with a real time transport packet header extension; optionally determining 603 that the second media stream is not associated with a real time transport packet header extension; parsing 604 the information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and using 605 the real time transport packet header extension of the first media stream for processing the second media stream. [0137] Each of the steps can be implemented by a respective module of a computer system.
[0138] An apparatus according to an embodiment is shown in Fig. 7. The apparatus is, for example, the server 102. The apparatus comprises means for receiving 111 at least a first media stream and a second media stream of a coded volumetric video; means for forming 112 a session description containing information about the at least the first media stream and the second media stream; means for associating 113 the first media stream a real time transport packet header extension; means for leaving 114 the second media stream without a real time transport packet header extension; and means for providing 115 information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream. The means comprises at least one processor, and a memory including a computer program code, wherein the processor may further comprise processor circuitry. The memory and the computer program code are configured to, with the at least one processor, cause the apparatus to perform the method of Fig. 5 according to various embodiments.
[0139] An apparatus according to another embodiment is shown in Fig. 8. The apparatus is, for example, the client 101. The apparatus comprises means for receiving 121 a session description containing information about at least a first media stream and a second media stream of a coded volumetric video; means for determining 122 that the first media stream is associated with a real time transport packet header extension; means for determining 123 that the second media stream is not associated with a real time transport packet header extension; means for parsing 124 the information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and means for using 125 the a real time transport packet header extension of the first media stream for processing the second media stream. The means comprises at least one processor, and a memory including a computer program code, wherein the processor may further comprise processor circuitry. The memory and the computer program code are configured to, with the at least one processor, cause the apparatus to perform the method of Fig. 6 according to various embodiments.
[0140] An example of an apparatus is disclosed with reference to Fig. 9. Fig. 9 shows a block diagram of a video coding system according to an example embodiment as a schematic block diagram of an electronic device 50, which may incorporate a codec. In some embodiments the electronic device may comprise an encoder or a decoder. The electronic device 50 may for example be a mobile terminal or a user equipment of a wireless communication system or a camera device. The electronic device 50 may be also comprised at a local or a remote server or a graphics processing unit of a computer. The device may be also comprised as part of a head-mounted display device. The apparatus 50 may comprise a display 32 in the form of a liquid crystal display. In other embodiments of the invention the display may be any suitable display technology suitable to display an image or video. The apparatus 50 may further comprise a keypad 34. In other embodiments of the invention any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display. The apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input. The apparatus 50 may further comprise an audio output device which in embodiments of the invention may be any one of: an earpiece 38, speaker, or an analogue audio or digital audio output connection. The apparatus 50 may also comprise a battery (or in other embodiments of the invention the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera 42 capable of recording or capturing images and/or video. The camera 42 may be a multi-lens camera system having at least two camera sensors. The camera is capable of recording or detecting individual frames which are then passed to the codec 54 or the controller for processing. The apparatus may receive the video and/or image data for processing from another device prior to transmission and/or storage.
[0141] The apparatus 50 may comprise a controller 56 or processor for controlling the apparatus 50. The apparatus or the controller 56 may comprise one or more processors or processor circuitry and be connected to memory 58 which may store data in the form of image, video and/or audio data, and/or may also store instructions for implementation on the controller 56 or to be executed by the processors or the processor circuitry. The controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and decoding of image, video and/or audio data or assisting in coding and decoding carried out by the controller.
[0142] The apparatus 50 may further comprise a card reader 48 and a smart card 46, for example a UICC (Universal Integrated Circuit Card) and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network. The apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and for receiving radio frequency signals from other apparatus(es). The apparatus may comprise one or more wired interfaces configured to transmit and/or receive data over a wired connection, for example an electrical cable or an optical fiber connection.
[0143] The various embodiments can be implemented with the help of computer program code that resides in a memory and causes the relevant apparatuses to carry out the method. For example, a device may comprise circuitry and electronics for handling, receiving, and transmitting data, computer program code in a memory, and a processor that, when running the computer program code, causes the device to carry out the features of an embodiment. Yet further, a network device like a server may comprise circuitry and electronics for handling, receiving, and transmitting data, computer program code in a memory, and a processor that, when running the computer program code, causes the network device to carry out the features of various embodiments.
[0144] If desired, the different functions discussed herein may be performed in a different order and/or concurrently with other. Furthermore, if desired, one or more of the above-described functions and embodiments may be optional or may be combined. [0145] Although various aspects of the embodiments are set out in the independent claims, other aspects comprise other combinations of features from the described embodiments and/or the dependent claims with the features of the independent claims, and not solely the combinations explicitly set out in the claims.
[0146] It is also noted herein that while the above describes example embodiments, these descriptions should not be viewed in a limiting sense. Rather, there are several variations and modifications, which may be made without departing from the scope of the present disclosure as, defined in the appended claims.

Claims

Claims:
1. An apparatus comprising: means for forming a session description containing information about at least a first media stream and a second media stream; means for receiving at least the first media stream and the second media stream; means for associating the first media stream a real time transport packet header extension; and means for providing reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
2. The apparatus according to claim 1 , wherein the means for providing the reuse information are configured to include the reuse information in extension attributes of a real time transport packet header extension mapping attribute.
3. The apparatus according to claim 1 or 2, wherein the means for providing the reuse information are configured to express the reuse information as a=hdrext- reuse : <value> <mid-li st>, in which <value> indicates which other media the header extension is applicable to, and <mid-li st> contains identifiers of the other media.
4. The apparatus according to claim 1 or 2, wherein the means for providing the reuse information are configured to express the reuse information implicitly by indicating that the first media stream and the second media stream belong to a same bundle.
5. The apparatus according to any of the claims 1 to 3, wherein the means for providing the reuse information are configured to express one or more additional parameters related to the first media stream and the second media stream.
6. The apparatus according to any of the claims 1 to 3, wherein the reuse information comprises an identifier of the header extension of the first media stream, information of one or more other media streams the header extension is applicable to, and zero, one or more attributes providing further information for the first media stream and the one or more other media streams.
7. The apparatus according to any one of the claims 1 to 6 comprising means for providing the reuse information in the real time transport packet header extension.
8. The apparatus according to claim 7 comprising means for providing identification of the second media stream in the reuse information.
9. The apparatus according to any one of the claims 1 to 8 comprising means for attaching the real time transport packet header extension to multiple streams at different frames, wherein the reuse information indicating that the real time transport packet header extension is flexibly present in different streams.
10. The apparatus according to any one of the claims 1 to 9, wherein the first media stream and the second media stream are streams of different components of a volumetric video.
11. The apparatus according to claim 10, wherein the first media stream media is a video stream and the second media stream is audio stream, or the first media stream and the second media stream are streams of different subpictures of a two-dimensional video, or the first media stream and the second media stream are streams of different layers of the video.
12. A method, comprising: forming a session description containing information about at least a first media stream and a second media stream; receiving at least the first media stream and the second media stream; associating the first media stream a real time transport packet header extension; and providing reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
13. An apparatus comprising: means for receiving a session description containing information about at least a first media stream and a second media stream; means for determining that the first media stream is associated with a real time transport packet header extension; means for parsing a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and means for using the real time transport packet header extension of the first media stream for processing the second media stream.
14. The apparatus according to claim 13, wherein the means for parsing the reuse information are configured to deduce the reuse information implicitly by determining that the first media stream and the second media stream belong to a same bundle.
15. The apparatus according to claim 13 to 14, wherein the means for parsing the reuse information are configured to obtain one or more additional parameters related to the first media stream and the second media stream.
16. The apparatus according to any one of the claims 13 to 15, wherein the reuse information comprises an identifier of the header extension of the first media stream, information of one or more other media streams the header extension is applicable to, and zero, one or more attributes providing further information for the first media stream and the one or more other media streams.
17. The apparatus according to any one of the claims 13 to 16 comprising one or more of the following: means for parsing the reuse information from the real time transport packet header extension; means for obtaining identification of the second media stream from the reuse information; means for parsing the reuse information from extension attributes in the real time transport packet header extension.
18. A method comprising: receiving a session description containing information about at least a first media stream and a second media stream; determining that the first media stream is associated with a real time transport packet header extension; parsing a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and using the real time transport packet header extension of the first media stream for processing the second media stream.
19. An apparatus comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following: forming a session description containing information about at least a first media stream and a second media stream; receiving at least the first media stream and the second media stream; associating the first media stream a real time transport packet header extension; leaving the second media stream without a real time transport packet header extension; and providing reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream.
20. An apparatus comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following: receiving a session description containing information about at least a first media stream and a second media stream; determining that the first media stream is associated with a real time transport packet header extension; parsing a reuse information indicating that the real time transport packet header extension of the first media stream can be utilized for processing of the second media stream; and using the real time transport packet header extension of the first media stream for processing the second media stream.
EP24774313.1A 2023-03-22 2024-02-20 A method, an apparatus and a computer program product for video coding Pending EP4684531A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
FI20235339 2023-03-22
PCT/FI2024/050066 WO2024194520A1 (en) 2023-03-22 2024-02-20 A method, an apparatus and a computer program product for video coding

Publications (1)

Publication Number Publication Date
EP4684531A1 true EP4684531A1 (en) 2026-01-28

Family

ID=92840979

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24774313.1A Pending EP4684531A1 (en) 2023-03-22 2024-02-20 A method, an apparatus and a computer program product for video coding

Country Status (2)

Country Link
EP (1) EP4684531A1 (en)
WO (1) WO2024194520A1 (en)

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
GB2563439B (en) * 2017-06-16 2022-02-16 Canon Kk Methods, devices, and computer programs for improving streaming of portions of media data

Also Published As

Publication number Publication date
WO2024194520A1 (en) 2024-09-26

Similar Documents

Publication Publication Date Title
TWI846795B (en) Multiple decoder interface for streamed media data
TWI423679B (en) Scalable video coding and decoding
US20220407899A1 (en) Real-time augmented reality communication session
EP4088476B1 (en) Multiple decoder interface for streamed media data
WO2023073283A1 (en) A method, an apparatus and a computer program product for video encoding and video decoding
EP4520031A1 (en) 5g support for webrtc
US20240275826A1 (en) Network rendering and transcoding of augmented reality data
US11863767B2 (en) Transporting HEIF-formatted images over real-time transport protocol
US12581098B2 (en) Transporting HEIF-formatted images over real-time transport protocol
WO2023062271A1 (en) A method, an apparatus and a computer program product for video coding
WO2024194520A1 (en) A method, an apparatus and a computer program product for video coding
US12375715B2 (en) Method, an apparatus and a computer program product for streaming volumetric video content
US12375634B2 (en) Method, an apparatus and a computer program product for spatial computing service session description for volumetric extended reality conversation
CN117099375A (en) Transfer HEIF-formatted images via real-time transfer protocol
EP4284000A1 (en) An apparatus, a method and a computer program for volumetric video
WO2024069045A9 (en) An apparatus and a method for processing volumetric video content
WO2023175234A1 (en) A method, an apparatus and a computer program product for streaming volumetric video
EP4442000A1 (en) A method, an apparatus and a computer program product for video encoding and video decoding
WO2023161556A1 (en) A method, an apparatus and a computer program product for video encoding and video decoding
WO2022266457A1 (en) Real-time augmented reality communication session
WO2024173339A1 (en) Network rendering and transcoding of augmented reality data

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251022

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR