EP4595447A1 - Split-rendering configuration for multimedia immersion and interaction data in a wireless communication system - Google Patents
Split-rendering configuration for multimedia immersion and interaction data in a wireless communication systemInfo
- Publication number
- EP4595447A1 EP4595447A1 EP23727506.0A EP23727506A EP4595447A1 EP 4595447 A1 EP4595447 A1 EP 4595447A1 EP 23727506 A EP23727506 A EP 23727506A EP 4595447 A1 EP4595447 A1 EP 4595447A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- video
- data
- media
- user
- configuration
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/235—Processing of additional data, e.g. scrambling of additional data or processing content descriptors
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/236—Assembling of a multiplex stream, e.g. transport stream, by combining a video stream with other content or additional data, e.g. inserting a URL [Uniform Resource Locator] into a video stream, multiplexing software data into a video stream; Remultiplexing of multiplex streams; Insertion of stuffing bits into the multiplex stream, e.g. to obtain a constant bit-rate; Assembling of a packetised elementary stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/236—Assembling of a multiplex stream, e.g. transport stream, by combining a video stream with other content or additional data, e.g. inserting a URL [Uniform Resource Locator] into a video stream, multiplexing software data into a video stream; Remultiplexing of multiplex streams; Insertion of stuffing bits into the multiplex stream, e.g. to obtain a constant bit-rate; Assembling of a packetised elementary stream
- H04N21/23614—Multiplexing of additional data and video streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/434—Disassembling of a multiplex stream, e.g. demultiplexing audio and video streams, extraction of additional data from a video stream; Remultiplexing of multiplex streams; Extraction or processing of SI; Disassembling of packetised elementary stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/434—Disassembling of a multiplex stream, e.g. demultiplexing audio and video streams, extraction of additional data from a video stream; Remultiplexing of multiplex streams; Extraction or processing of SI; Disassembling of packetised elementary stream
- H04N21/4348—Demultiplexing of additional data and video streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/60—Network structure or processes for video distribution between server and client or between remote clients; Control signalling between clients, server and network components; Transmission of management data between server and client, e.g. sending from server to client commands for recording incoming content stream; Communication details between server and client
- H04N21/65—Transmission of management data between client and server
- H04N21/658—Transmission by the client directed to the server
- H04N21/6587—Control parameters, e.g. trick play commands, viewpoint selection
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/81—Monomedia components thereof
- H04N21/816—Monomedia components thereof involving special video data, e.g 3D video
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/85—Assembly of content; Generation of multimedia applications
- H04N21/854—Content authoring
- H04N21/85406—Content authoring involving a specific file format, e.g. MP4 format
Definitions
- XR is used as an umbrella term for different types of realities, according to 3GPP Technical Report TR 26.928 (v17.0.0 – Apr 2022). These types of realities include, Virtual Reality (VR), Augmented Reality (AR) and Mixed Reality (MR).
- VR is a rendered version of a delivered visual and audio scene.
- the rendering is in this case designed to mimic the visual and audio sensory stimuli of the real world as naturally as possible to an observer or user as they move within the limits defined by the application.
- Virtual reality usually, but not necessarily, requires a user to wear a head mounted display (HMD), to completely replace the user's field of view with a simulated visual component, and to wear headphones, to provide the user with the accompanying audio.
- HMD head mounted display
- Some form of head and motion tracking of the user in VR is usually also necessary SMM920220294-GR-NP to allow the simulated visual and audio components to be updated to ensure that, from the user's perspective, items and sound sources remain consistent with the user's movements.
- additional means to interact with the virtual reality simulation may be provided but are not strictly necessary.
- AR is when a user is provided with additional information or artificially generated items, or content overlaid upon their current environment. Such additional information or content will usually be visual and/or audible and their observation of their current environment may be direct, with no intermediate sensing, processing, and rendering, or indirect, where their perception of their environment is relayed via sensors and may be enhanced or processed.
- MR is an advanced form of AR where some virtual elements are inserted into the physical scene with the intent to provide the illusion that these elements are part of the real scene.
- XR refers to all real-and-virtual combined environments and human-machine interactions generated by computer technology and wearables. It includes representative forms such as AR, MR and VR and the areas interpolated among them.
- the levels of virtuality range from partially sensory inputs to fully immersive VR.
- a key aspect of XR is the extension of human experiences especially relating to the senses of existence (represented by VR) and the acquisition of cognition (represented by AR).
- VR senses of existence
- AR acquisition of cognition
- CG Cloud Gaming
- the data such applications carry and leverage to generate the cyber-physical immersiveness illusion has been categorized into a number of classes as set out in 3GPP Technical Document S4-221557 (Nov 2022), or alternatively, Technical Report TR 26.926 (v1.1.0 – Feb 2022).
- the formats associated with this data class describe the physical and hardware capabilities of an end user equipment (UE) and/or glass device.
- Some examples in this sense are camera sub-system capabilities and camera configuration (e.g., focal length, available zoom, and depth calibration information, pose SMM920220294-GR-NP reference of the main camera etc.), projection formats (e.g., cubemap, equirectangular, fisheye, stereographic etc.).
- the device capability data is usually static and available before the establishment of a session, hence its transfer and transport over a network is not of high concern as it can be embedded in typical session configuration procedures and protocols, such as Session Initiation Protocol (SIP) and/or Session Description Protocol (SDP).
- SIP Session Initiation Protocol
- SDP Session Description Protocol
- the device capability data is as such not real-time sensitive and has no real-time transport requirements.
- the data describes the space and/or the object content of a view.
- this data can be a scene description used to detail the 3D composition of space anchoring 2D and 3D objects within a scene (e.g., as a tree or graph structure usually of glTF2.0 or JSON syntax).
- Another possible representation is of a spatial description used for spatial computing and mapping of the real-world to its virtual counterpart or vice versa.
- this data type may contain 3D model descriptors of objects and their attributes formatted for instance as meshes (i.e., sets of vertices, edges and faces), or point cloud data formatted under PoLYgon (PLY) syntax to be consumed by the visual presentation devices, i.e., the UEs.
- Other data types may represent dynamic world graph representations whereby selected trackables (e.g., geo-cached AR/QR codes, geo-trackables like physical objects located at a specified world position, dynamic physical objects like buses, subways, etc.) enter and leave the scene perspective of the world dynamically and need to be conveyed in real-time to an AR runtime.
- the media description class of data may be of large size (i.e., often even more than 10 MBytes) and it may be updated with low frequency (within 10s of seconds regime) under various event triggers (e.g., user viewport change, new object entering the scene, old object exiting the scene, scene change and/or update etc.).
- the media description data may be real-time sensitive as it is involved in completing the display of the virtual renderings to a presentation device such as a UE, and as a result may benefit of real-time transport over a network.
- this data type contains user spatial interaction information such as: user viewport description (i.e., an encoding of azimuth, elevation, tilt, and associated ranges of motion describing the projection of the user view to a target display); user field of view (FoV) (i.e., the extent of the visible world from the viewer perspective usually described in angular domain, e.g., radians/degrees, over vertical and horizontal planes); user pose/orientation tracking data (i.e., micro- /nanosecond timestamped 3D vector for position and quaternion representation for SMM920220294-GR-NP orientation describing up to 6DoF); user gesture tracking data (i.e., an array of hands tracked each consisting of an array of hand joint locations relative to a base space); [0013] user body tracking data (e.g., a BioVision Hierarchical BVH encoding of the body and body segments movements); user facial expression/eye movement tracking data (e.g.
- user viewport description i.e.
- the interaction and immersion class data has certain characteristics. These characteristics include: low data footprint ranging usually from 32 Bytes up to around hundreds of Bytes per message with no established codecs for compression of data sources; high sampling rates varying between the video FPS frequency, e.g., 60 Hz up to 250 Hz, and in some cases wherein sample aggregation is not performed raw sample reports may be transmitted even at 1000 Hz sampling frequency; data can trigger a response with low-latency requirements (e.g., up to 50 milliseconds end-to-end from the interaction to the response as perceived by the user); can be synchronized to other media streams (e.g., video or audio media stream); can be synchronized to other interaction data (e.g., pose information may be synchronized with user actions, or alternatively, object actions); reliability is optional as determined by individual application requirements’ (e.g., in split-rendering scenarios servers may predict future pose estimates based on available pose information and hence high reliability below 10 ⁇ (-3) error rate is not necessary);
- the interaction and immersion metadata class is therefore real-time sensitive and requires real-time transport over a network on par with existent solutions for established media flows pertaining to video, or alternatively, audio codecs.
- the format and syntax of such information flows is often application, platform and/or HW dependent and in contradiction to well-established media formats and codecs (e.g., audio or video codecs), no mainstream encodings, syntax and semantics are well established.
- no mainstream transport specific solutions have been established yet as such information flows and associated data formats evolve rapidly. The latter fact requires fast adaptation to new versions at a higher rate than typical conventional media codecs development cycles.
- the media description and interaction data benefitting real-time transmission and synchronization may be in some implementations transmitted based on at least three different options based on existing technologies. These technologies are: WebRTC data channel based on the Stream Control Transmission Protocol (SCTP); RTP header extension embedding the metadata information in-band in the RTP transport (for example, US Patent Application #63/420,885); and new RTP payload format generically dedicated to the transport of real-time metadata (for example, US Patent Application #63/478,932).
- SCTP Stream Control Transmission Protocol
- RTP header extension embedding the metadata information in-band in the RTP transport for example, US Patent Application #63/420,885
- new RTP payload format generically dedicated to the transport of real-time metadata for example, US Patent Application #63/478,932).
- a potential solution for the first technology comprises the SCTP data channel of the WebRTC being used to carry interaction metadata.
- a generic data channel payload format for timed metadata including a timestamp is added to the chunk user data section of the SCTP data channel.
- a potential solution (discussed in 3GPP Tdoc S4-221555, or alternatively, US Patent Application #63/420,885) comprises a RTP SMM920220294-GR-NP header extension designed to carry interaction metadata of limited size while its associated media content is carried in the RTP payload.
- Support for a single metadata type or multiple metadata types can be carried in the proposed header extension to allow for scalability and flexibility.
- This approach has the advantage that the transported metadata is time-synchronized to the media data.
- all the robustness and timing mechanisms provided by RTP are included (e.g., synchronization, jitter, congestion control support, FEC mechanisms etc.).
- sending interaction metadata in a separate RTP stream is based on defining a new RTP payload format for interaction metadata whereby a generic RTP metadata payload dedicated to the transport of different interaction and immersion metadata is formulated.
- the advantage of this approach is that it enables the usage of all RTP mechanisms (timing, synchronization, jitter management support as well as FEC robustness etc) while providing a generic format that can cover all types of interaction metadata.
- This requires carry over in IETF to become a universal transport standard.
- definition of a new payload format typically takes at least two years in IETF meaning that for 3GPP, or alike, networked systems, the developed format would at the earliest be useful at the end of the 3GPP Release 19 cycle, or alternatively, in the medium-term.
- an apparatus for wireless communication comprising a processor; and a memory coupled with the processor, the processor configured to cause the apparatus to: determine a media configuration for split-rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating an encoded video stream of the video media flow comprises, as non-video coded metadata, one or more data units of multimedia immersion and interaction data; signal the media configuration to a second apparatus; establish, with the second apparatus, based at least in part on the media configuration, a multimedia split-rendering content delivery session comprising the video media flow; and use, for split-rendering the video media flow, the one or more data units of multimedia immersion and interaction data.
- an apparatus for wireless communication comprising a processor; and a memory coupled with the processor, the processor configured to cause the apparatus to: receive, from a first apparatus, a media configuration for split-rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating an encoded video stream of the video media flow comprises, as non-video coded metadata, one or more data units of multimedia immersion and interaction data; configure the apparatus, using the media configuration, to receive the video media flow; decode the encoded video stream of the video media flow, wherein the decoding comprises extracting the non-video coded metadata; and consume, from the non-video coded metadata, the one or more data units of multimedia immersion and interaction data.
- a method of wireless communication comprising: determining a media configuration for split-rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating an encoded video stream of the video media flow comprises, as non-video coded metadata, one or more data units of multimedia immersion and interaction data; signaling the media configuration to a second apparatus; establishing, with the second apparatus, based at SMM920220294-GR-NP least in part on the media configuration, a multimedia split-rendering content delivery session comprising the video media flow; and using, for split-rendering the video media flow, the one or more data units of multimedia immersion and interaction data.
- a method for wireless communication comprising: receiving, from a first apparatus, a media configuration for split-rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating an encoded video stream of the video media flow comprises, as non-video coded metadata, one or more data units of multimedia immersion and interaction data; configuring the apparatus, using the media configuration, to receive the video media flow; decoding the encoded video stream of the video media flow, wherein the decoding comprises extracting the non-video coded metadata; and consuming, from the non-video coded metadata, the one or more data units of multimedia immersion and interaction data.
- Figure 1 illustrates an embodiment of a wireless communication system
- Figure 2 illustrates an embodiment of a user equipment apparatus
- Figure 3 illustrates an embodiment of a network node
- Figure 4 illustrates an RTP and RTCP protocol stack over IP networks
- Figure 5 illustrates a WebRTC (SRTP) protocol stack over IP networks
- Figure 6 illustrates an RTP packet format and header information
- Figure 7 illustrates an SRTP packet format and header information
- Figure 8 illustrates an RTP/SRTP header extension format and syntax
- Figure 9 illustrates a simplified block diagram of a generic video codec performing spatial and temporal compression of a video source
- Figure 10 illustrates a video coded elementary stream and corresponding plurality of NAL units
- Figure 11 illustrates an embodiment of a method of wireless communication in a wireless communication system
- Figure 12 illustrates an alternative
- aspects of this disclosure may be embodied as a system, apparatus, method, or program product. Accordingly, arrangements described herein may be implemented in an entirely hardware form, an entirely software form (including firmware, resident software, micro-code, etc.) or a form combining software and hardware aspects.
- the disclosed methods and apparatus may be implemented as a hardware circuit comprising custom very-large-scale integration (“VLSI”) circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components.
- VLSI very-large-scale integration
- the disclosed methods and apparatus may also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, or the like.
- the disclosed methods and apparatus may include one or more physical or logical blocks of executable code which may, for instance, be organized as an object, procedure, or function.
- the methods and apparatus may take the form of a program product embodied in one or more computer readable storage devices storing machine readable code, computer readable code, and/or program code, referred hereafter as code.
- the storage devices may be tangible, non-transitory, and/or non-transmission.
- the SMM920220294-GR-NP storage devices may not embody signals. In certain arrangements, the storage devices only employ signals for accessing code.
- Any combination of one or more computer readable medium may be utilized.
- the computer readable medium may be a computer readable storage medium.
- the computer readable storage medium may be a storage device storing the code.
- the storage device may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, holographic, micromechanical, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
- a storage device More specific examples (a non-exhaustive list) of the storage device would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random-access memory (“RAM”), a read-only memory (“ROM”), an erasable programmable read-only memory (“EPROM” or Flash memory), a portable compact disc read-only memory (“CD-ROM”), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
- a computer readable storage medium may be any tangible medium that can contain, or store, a program for use by or in connection with an instruction execution system, apparatus, or device.
- references throughout this specification to an example of a particular method or apparatus, or similar language means that a particular feature, structure, or characteristic described in connection with that example is included in at least one implementation of the method and apparatus described herein.
- reference to features of an example of a particular method or apparatus, or similar language may, but do not necessarily, all refer to the same example, but mean “one or more but not all examples” unless expressly specified otherwise.
- a list with a conjunction of “and/or” includes any single item in the list or a combination of items in the list.
- a list of A, B and/or C includes only A, only B, only C, a combination of A and B, a combination of B and C, a combination of A and C or a combination of A, B and C.
- a list using the terminology “one or more of” includes any single item in the list or a combination of items in the list.
- one or more of A, B and C includes only A, only B, only SMM920220294-GR-NP C, a combination of A and B, a combination of B and C, a combination of A and C or a combination of A, B and C.
- a list using the terminology “one of” includes one, and only one, of any single item in the list.
- “one of A, B and C” includes only A, only B or only C and excludes combinations of A, B and C.
- a member selected from the group consisting of A, B, and C includes one and only one of A, B, or C, and excludes combinations of A, B, and C.”
- a member selected from the group consisting of A, B, and C and combinations thereof includes only A, only B, only C, a combination of A and B, a combination of B and C, a combination of A and C or a combination of A, B and C.
- each block of the schematic flowchart diagrams and/or schematic block diagrams, and combinations of blocks in the schematic flowchart diagrams and/or schematic block diagrams can be implemented by code.
- This code may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions/acts specified in the schematic flowchart diagrams and/or schematic block diagrams.
- the code may also be stored in a storage device that can direct a computer, other programmable data processing apparatus, or other devices to function in a particular manner, such that the instructions stored in the storage device produce an article of SMM920220294-GR-NP manufacture including instructions which implement the function/act specified in the schematic flowchart diagrams and/or schematic block diagrams.
- the code may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer implemented process such that the code which executes on the computer or other programmable apparatus provides processes for implementing the functions/acts specified in the schematic flowchart diagrams and/or schematic block diagram.
- each block in the schematic flowchart diagrams and/or schematic block diagrams may represent a module, segment, or portion of code, which includes one or more executable instructions of the code for implementing the specified logical function(s).
- the functions noted in the block may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.
- Figure 1 depicts an embodiment of a wireless communication system 100 for a split-rendering configuration for multimedia immersion and interaction data in a wireless communication system.
- the wireless communication system 100 includes remote units 102 and network units 104. Even though a specific number of remote units 102 and network units 104 are depicted in Figure 1, one of skill in the art will recognize that any number of remote units 102 and network units 104 may be included in the wireless communication system 100.
- the wireless communication system may comprise a wireless communication network and at least one wireless communication device.
- the wireless communication device is typically a 3GPP User Equipment (UE).
- the wireless communication network may comprise at least one network node.
- the network node may be a network unit.
- SMM920220294-GR-NP [0044]
- the remote units 102 may include computing devices, such as desktop computers, laptop computers, personal digital assistants (“PDAs”), tablet computers, smart phones, smart televisions (e.g., televisions connected to the Internet), set-top boxes, game consoles, security systems (including security cameras), vehicle on- board computers, network devices (e.g., routers, switches, modems), aerial vehicles, drones, or the like.
- the remote units 102 include wearable devices, such as smart watches, fitness bands, optical head-mounted displays, or the like. Moreover, the remote units 102 may be referred to as subscriber units, mobiles, mobile stations, users, terminals, mobile terminals, fixed terminals, subscriber stations, UE, user terminals, a device, or by other terminology used in the art.
- the remote units 102 may communicate directly with one or more of the network units 104 via UL communication signals. In certain embodiments, the remote units 102 may communicate directly with other remote units 102 via sidelink communication.
- the network units 104 may be distributed over a geographic region.
- a network unit 104 may also be referred to as an access point, an access terminal, a base, a base station, a Node-B, an eNB, a gNB, a Home Node-B, a relay node, a device, a core network, an aerial server, a radio access node, an AP, NR, a network entity, an Access and Mobility Management Function (“AMF”), a Unified Data Management Function (“UDM”), a Unified Data Repository (“UDR”), a UDM/UDR, a Policy Control Function (“PCF”), a Radio Access Network (“RAN”), an Network Slice Selection Function (“NSSF”), an operations, administration, and management (“OAM”), a session management function (“SMF”), a user plane function (“UPF”), an application function, an authentication server function (“AUSF”), security anchor functionality (“SEAF”), trusted non-3GPP gateway function (“TNGF”), an application function, a service enabler architecture layer (“SEAL”) function, a
- AMF
- the network units 104 are generally part of a radio access network that includes one or more controllers communicably coupled to one or more corresponding network units 104.
- the radio access network is generally communicably coupled to one or more core networks, which may be coupled to other networks, like the Internet and public switched telephone networks, among other networks.
- the wireless communication system 100 is compliant with New Radio (NR) protocols standardized in 3GPP, wherein the network unit 104 transmits using an Orthogonal Frequency Division Multiplexing (“OFDM”) modulation scheme on the downlink (DL) and the remote units 102 transmit on the uplink (UL) using a Single Carrier Frequency Division Multiple Access (“SC-FDMA”) scheme or an OFDM scheme.
- OFDM Orthogonal Frequency Division Multiplexing
- SC-FDMA Single Carrier Frequency Division Multiple Access
- the wireless communication system 100 may implement some other open or proprietary communication protocol, for example, WiMAX, IEEE 802.11 variants, GSM, GPRS, UMTS, LTE variants, CDMA2000, Bluetooth®, ZigBee, Sigfox, LoraWAN among other protocols.
- the present disclosure is not intended to be limited to the implementation of any particular wireless communication system architecture or protocol.
- the network units 104 may serve a number of remote units 102 within a serving area, for example, a cell or a cell sector via a wireless communication link.
- the network units 104 transmit DL communication signals to serve the remote units 102 in the time, frequency, and/or spatial domain.
- Figure 2 depicts a user equipment apparatus 200 that may be used for implementing the methods described herein.
- the user equipment apparatus 200 is used to implement one or more of the solutions described herein.
- the user equipment apparatus 200 is in accordance with one or more of the user equipment apparatuses described in embodiments herein.
- the user equipment apparatus 200 may comprise a UE 102 or a UE 1310 of Figure 13a, comprising application client 1311, MSH 1312, split rendering client 1313, for instance.
- the user equipment apparatus 200 includes a processor 205, a memory 210, an input device 215, an output device 220, and a transceiver 225.
- the input device 215 and the output device 220 may be combined into a single device, such as a touchscreen. In some implementations, the user equipment apparatus 200 does not include any input device 215 and/or output device 220.
- the user equipment apparatus 200 may include one or more of: the processor 205, the memory 210, and the transceiver 225, and may not include the input device 215 and/or the output device 220.
- the transceiver 225 includes at least one transmitter 230 and at least one receiver 235.
- the transceiver 225 may communicate with one or more cells (or SMM920220294-GR-NP wireless coverage areas) supported by one or more base units.
- the transceiver 225 may be operable on unlicensed spectrum.
- the transceiver 225 may include multiple UE panels supporting one or more beams. Additionally, the transceiver 225 may support at least one network interface 240 and/or application interface 245.
- the application interface(s) 245 may support one or more APIs.
- the network interface(s) 240 may support 3GPP reference points, such as Uu, N1, PC5, etc. Other network interfaces 240 may be supported, as understood by one of ordinary skill in the art.
- the processor 205 may include any known controller capable of executing computer-readable instructions and/or capable of performing logical operations.
- the processor 205 may be a microcontroller, a microprocessor, a central processing unit (“CPU”), a graphics processing unit (“GPU”), an auxiliary processing unit, a field programmable gate array (“FPGA”), or similar programmable controller.
- the processor 205 may execute instructions stored in the memory 210 to perform the methods and routines described herein.
- the processor 205 is communicatively coupled to the memory 210, the input device 215, the output device 220, and the transceiver 225.
- the processor 205 may control the user equipment apparatus 200 to implement the user equipment apparatus behaviors described herein.
- the processor 205 may include an application processor (also known as “main processor”) which manages application-domain and operating system (“OS”) functions and a baseband processor (also known as “baseband radio processor”) which manages radio functions.
- the memory 210 may be a computer readable storage medium.
- the memory 210 may include volatile computer storage media.
- the memory 210 may include a RAM, including dynamic RAM (“DRAM”), synchronous dynamic RAM (“SDRAM”), and/or static RAM (“SRAM”).
- DRAM dynamic RAM
- SDRAM synchronous dynamic RAM
- SRAM static RAM
- the memory 210 may include non-volatile computer storage media.
- the memory 210 may include a hard disk drive, a flash memory, or any other suitable non-volatile computer storage device.
- the memory 210 may include both volatile and non-volatile computer storage media.
- the memory 210 may store data related to implement a traffic category field as described herein.
- the memory 210 may also store program code and related data, such as an operating system or other controller algorithms operating on the apparatus 200.
- the input device 215 may include any known computer input device including a touch panel, a button, a keyboard, a stylus, a microphone, or the like.
- the input device 215 may be integrated with the output device 220, for example, as a touchscreen or similar touch-sensitive display.
- the input device 215 may include a touchscreen such SMM920220294-GR-NP that text may be input using a virtual keyboard displayed on the touchscreen and/or by handwriting on the touchscreen.
- the input device 215 may include two or more different devices, such as a keyboard and a touch panel.
- the output device 220 may be designed to output visual, audible, and/or haptic signals.
- the output device 220 may include an electronically controllable display or display device capable of outputting visual data to a user.
- the output device 220 may include, but is not limited to, a Liquid Crystal Display (“LCD”), a Light- Emitting Diode (“LED”) display, an Organic LED (“OLED”) display, a projector, or similar display device capable of outputting images, text, or the like to a user.
- the output device 220 may include a wearable display separate from, but communicatively coupled to, the rest of the user equipment apparatus 200, such as a smart watch, smart glasses, a heads-up display, or the like.
- the output device 220 may be a component of a smart phone, a personal digital assistant, a television, a table computer, a notebook (laptop) computer, a personal computer, a vehicle dashboard, or the like.
- the output device 220 may include one or more speakers for producing sound.
- the output device 220 may produce an audible alert or notification (e.g., a beep or chime).
- the output device 220 may include one or more haptic devices for producing vibrations, motion, or other haptic feedback. All, or portions, of the output device 220 may be integrated with the input device 215.
- the input device 215 and output device 220 may form a touchscreen or similar touch-sensitive display.
- the output device 220 may be located near the input device 215.
- the transceiver 225 communicates with one or more network functions of a mobile communication network via one or more access networks.
- the transceiver 225 operates under the control of the processor 205 to transmit messages, data, and other signals and also to receive messages, data, and other signals.
- the processor 205 may selectively activate the transceiver 225 (or portions thereof) at particular times in order to send and receive messages.
- the transceiver 225 includes at least one transmitter 230 and at least one receiver 235.
- the one or more transmitters 230 may be used to provide uplink communication signals to a base unit of a wireless communication network.
- the one or more receivers 235 may be used to receive downlink communication signals from the base unit. Although only one transmitter 230 and one receiver 235 are illustrated, the user equipment apparatus 200 may have any suitable number of transmitters 230 and receivers SMM920220294-GR-NP 235. Further, the transmitter(s) 230 and the receiver(s) 235 may be any suitable type of transmitters and receivers.
- the transceiver 225 may include a first transmitter/receiver pair used to communicate with a mobile communication network over licensed radio spectrum and a second transmitter/receiver pair used to communicate with a mobile communication network over unlicensed radio spectrum.
- the first transmitter/receiver pair may be used to communicate with a mobile communication network over licensed radio spectrum and the second transmitter/receiver pair used to communicate with a mobile communication network over unlicensed radio spectrum may be combined into a single transceiver unit, for example a single chip performing functions for use with both licensed and unlicensed radio spectrum.
- the first transmitter/receiver pair and the second transmitter/receiver pair may share one or more hardware components.
- certain transceivers 225, transmitters 230, and receivers 235 may be implemented as physically separate components that access a shared hardware resource and/or software resource, such as for example, the network interface 240.
- One or more transmitters 230 and/or one or more receivers 235 may be implemented and/or integrated into a single hardware component, such as a multi- transceiver chip, a system-on-a-chip, an Application-Specific Integrated Circuit (“ASIC”), or other type of hardware component.
- ASIC Application-Specific Integrated Circuit
- One or more transmitters 230 and/or one or more receivers 235 may be implemented and/or integrated into a multi-chip module.
- Other components such as the network interface 240 or other hardware components/circuits may be integrated with any number of transmitters 230 and/or receivers 235 into a single chip.
- the transmitters 230 and receivers 235 may be logically configured as a transceiver 225 that uses one more common control signals or as modular transmitters 230 and receivers 235 implemented in the same hardware chip or in a multi-chip module.
- Figure 3 depicts further details of the network node 300 that may be used for implementing the methods described herein.
- the network node 300 may be one implementation of an entity in the wireless communication network, e.g. in one or more of the wireless communication networks described herein.
- the network node 300 may comprise a network node such as a node of the edge network 1320 in Figure 13a (for instance the SRAF 1321 and configuration function 1321a and provisioning function 1322b; the RTC AS 1322 and signaling server 1322a and SRS 1322b) or may be a node of SMM920220294-GR-NP the data network 1330 such as ASP 1331.
- the network node 300 includes a processor 305, a memory 310, an input device 315, an output device 320, and a transceiver 325. [0063]
- the input device 315 and the output device 320 may be combined into a single device, such as a touchscreen.
- the network node 300 does not include any input device 315 and/or output device 320.
- the network node 300 may include one or more of: the processor 305, the memory 310, and the transceiver 325, and may not include the input device 315 and/or the output device 320.
- the transceiver 325 includes at least one transmitter 330 and at least one receiver 335.
- the transceiver 325 communicates with one or more remote units 200.
- the transceiver 325 may support at least one network interface 340 and/or application interface 345.
- the application interface(s) 345 may support one or more APIs.
- the network interface(s) 340 may support 3GPP reference points, such as Uu, N1, N2 and N3. Other network interfaces 340 may be supported, as understood by one of ordinary skill in the art.
- the processor 305 may include any known controller capable of executing computer-readable instructions and/or capable of performing logical operations.
- the processor 305 may be a microcontroller, a microprocessor, a CPU, a GPU, an auxiliary processing unit, a FPGA, or similar programmable controller.
- the processor 305 may execute instructions stored in the memory 310 to perform the methods and routines described herein.
- the processor 305 is communicatively coupled to the memory 310, the input device 315, the output device 320, and the transceiver 325.
- the memory 310 may be a computer readable storage medium.
- the memory 310 may include volatile computer storage media.
- the memory 310 may include a RAM, including dynamic RAM (“DRAM”), synchronous dynamic RAM (“SDRAM”), and/or static RAM (“SRAM”).
- the memory 310 may include non-volatile computer storage media.
- the memory 310 may include a hard disk drive, a flash memory, or any other suitable non-volatile computer storage device.
- the memory 310 may include both volatile and non-volatile computer storage media.
- the memory 310 may store data related to establishing a multipath unicast link and/or mobile operation.
- the memory 310 may store parameters, configurations, resource assignments, policies, and the like, as described herein.
- the memory 310 may also store program code and related data, such as an operating system or other controller algorithms operating on the network node 300.
- SMM920220294-GR-NP may include any known computer input device including a touch panel, a button, a keyboard, a stylus, a microphone, or the like.
- the input device 315 may be integrated with the output device 320, for example, as a touchscreen or similar touch-sensitive display.
- the input device 315 may include a touchscreen such that text may be input using a virtual keyboard displayed on the touchscreen and/or by handwriting on the touchscreen.
- the input device 315 may include two or more different devices, such as a keyboard and a touch panel.
- the output device 320 may be designed to output visual, audible, and/or haptic signals.
- the output device 320 may include an electronically controllable display or display device capable of outputting visual data to a user.
- the output device 320 may include, but is not limited to, an LCD display, an LED display, an OLED display, a projector, or similar display device capable of outputting images, text, or the like to a user.
- the output device 320 may include a wearable display separate from, but communicatively coupled to, the rest of the network node 300, such as a smart watch, smart glasses, a heads-up display, or the like.
- the output device 320 may be a component of a smart phone, a personal digital assistant, a television, a table computer, a notebook (laptop) computer, a personal computer, a vehicle dashboard, or the like.
- the output device 320 may include one or more speakers for producing sound.
- the output device 320 may produce an audible alert or notification (e.g., a beep or chime).
- the output device 320 may include one or more haptic devices for producing vibrations, motion, or other haptic feedback. All, or portions, of the output device 320 may be integrated with the input device 315.
- the input device 315 and output device 320 may form a touchscreen or similar touch-sensitive display.
- the output device 320 may be located near the input device 315.
- the transceiver 325 includes at least one transmitter 330 and at least one receiver 335.
- the one or more transmitters 330 may be used to communicate with the UE, as described herein.
- the one or more receivers 335 may be used to communicate with network functions in the PLMN and/or RAN, as described herein.
- the network node 300 may have any suitable number of transmitters 330 and receivers 335.
- the transmitter(s) 330 and the receiver(s) 335 may be any suitable type of transmitters and receivers.
- RTP Real-time Transport Protocol
- SMM920220294-GR-NP RFC 3550 titled “RTP: A Transport Protocol for Real-Time Applications”
- SRTP Secure Real-time Transport Protocol
- SRTP Secure Real-time Transport Protocol
- W3C W3C standard recommendation dated 06 March 2023 titled “WebRTC: Real-Time Communication in Browsers”
- RTP is a media codec agnostic network protocol with application-layer framing used to deliver multimedia (e.g., audio, video etc.) data in real-time over IP networks. It is used in conjunction with a sister protocol for control, i.e., Real-time Transport Control Protocol (RTCP), to provide end-to-end features such as jitter compensation, packet loss and out-of-order delivery detection, synchronization and source streams multiplexing.
- RTCP Real-time Transport Control Protocol
- Figure 4 illustrates an overview of the RTP and RTCP protocol stack.
- An IP layer 405 carries signaling from the media session data plane 410 and from the media session control plane 450.
- the data plane 410 stack comprises functions for a User Datagram Protocol (UDP) 412, RTP 416, RTCP 414, Media codecs 420 and quality control 422.
- the control plane 450 stack comprises functions for UDP 452, Transmission Control Protocol (TCP) 454, Session Initiation Protocol (SIP) 462 and Session Description Protocol (SDP) 464.
- UDP User Datagram Protocol
- TCP Transmission Control Protocol
- SIP Session Initiation Protocol
- SDP Session Description Protocol
- SRTP is a secured version of RTP, providing encryption (mainly by means of payload confidentiality), message authentication and integrity protection (by means of PDU, i.e., headers and payload, signing), as well as replay attack protection.
- the SRTP sister protocol is SRTCP. This provides the same functions to its RTCP counterpart.
- FIG. 7 illustrates a overview of a WebRTC (i.e., based on SRTP) protocol stack.
- an IP layer 505 carries signaling from the data plane 510 and the control plane 550.
- the data plane 510 stack comprises functions for UDP 512, Interactive Connectivity Establishment (ICE) 524, Datagram Transport Layer Security (DTLS) 526, SRTP 517, SRTCP 515, media codecs 520, Quality Control 522 and SCTP 528.
- ICE 574 SMM920220294-GR-NP may use the Session Traversal Utilities for NAT (STUN) protocol and Traversal Using Relays around NAT (TURN) to address real-time media content delivery across heterogeneous networks and NAT rules and firewalls.
- the SCTP 528 data plane is mainly dedicated as an application data channel and may be non-time critical, whereas the SRTP 517 based stack including elements of control, i.e., SRTCP 515, encoding, i.e., media codecs 520, and Quality of Service (QoS), i.e., Quality Control 522, is dedicated to time-critical transport.
- the control plane 550 is shown as comprising TCP 554, TLS 556, HTTP 558, SSE/XHR/other 568, XMPP/other 570, SDP 564, and SIP 562.
- the RTP and SRTP header information share the same format as illustrated, respectively, in Figures 6 and 7.
- Figure 6 illustrates an RTP packet 630
- Figure 7 illustrates an SRTP packet 760
- V 641, 761
- ‘X’ 634, 764 is 1 bit indicating that the standard fixed RTP/SRTP header will be followed by an RTP header extension usually associated with a particular data/profile that will carry more information about the data (e.g., the frame marking RTP header extension for video data, as described in the IETF working draft dated November 2021 titled “Frame Marking RTP Header Extension”, or generic RTP header extensions such as the RTP/SRTP extended protocol, as described in IETF standard RFC 6904 titled “Encryption of Header Extensions in the Secure Real-time Transport Protocol (SRTP)”.) [0080] ‘CC’ 636, 766, is 4 bits indicating number of contributing media sources (CSRC) that follow the fixed header.
- CSRC contributing media sources
- ‘M’ 638, 768 is 1 bit intended to mark information frame boundaries in the packet stream, whose behaviour is exactly specified by RTP profiles (e.g., H.264, H.265, H.266, AV1 etc.).
- ‘PT’ 640, 770 is 7 bits indicating the payload type, which in case of audio and video codec profiles may be dynamic and negotiated by means of SDP (e.g., 96 for H.264, 97 for H.265, 98 for AV1 etc.).
- the payload profiles are registered with IANA and rely on IETF profiles describing how the transmission of data is enclosed within the payload of an RTP PDU.
- Sequence number 642, 772 is 16 bits indicating the sequence number which increments by one with each RTP data packet sent over a session.
- ‘Timestamp’ 644, 774 is 32 bits indicating timestamp in ticks of the payload type clock reflecting the sampling instant of the first octet of the RTP data packet (associated for video stream with a video frame), whereas the first timestamp of the first RTP packet is selected at random.
- ‘Synchronization Source (SSRC) identifier’ 646, 776 is 32 bits field indicating a random identifier for the source of a stream of RTP packets forming a part of the same timing and sequence number space, such that a receiver may group packets based on synchronization source for playback.
- ‘Contributing Source (CSRC) identifier’ 648, 778 list of up to 16 CSRC items of 32 bits each given the amount of CSRC mixed by RTP mixers within the current payload as signalled by the CC bits. The list identifies the contributing sources for the payload contained in this packet given the SSRC identifiers of the contributing sources.
- RTP header extension 650, 780 is a variable length field present if the X bit 634, 764 is marked. The header extension is appended to the RTP fixed header information after the CSRC list 648, 778 if present.
- the RTP header extension 650, 780 is 32-bit aligned and formed of the following fields: A 16-bit extension identifier defined by a profile and usually negotiated and determined via the Session Description Protocol (SDP) signaling mechanism; a 16-bit length field describing the extension header length in 32-bits multiples excluding the first 32 bits corresponding to the 16 bits extension identifier and the 16 bits length fields itself; and a 32-bit aligned header extension raw data field formatted according to some RTP header extension identifier specified format.
- SDP Session Description Protocol
- the RTP header extension 650, 780 format and syntax are like the ones of SRTP.
- the format 800 and syntax is illustrated in Figure 8.
- RTP extension header 650, 780 may be appended to the fixed header information, as described in IETF standard RFC 3550 titled “RTP: A Transport Protocol for Real-Time Applications”.
- RTP and SRTP extensions to the base protocols exist to allow for multiple RTP header extensions 650, 780 of predetermined SMM920220294-GR-NP types to be appended to the fixed header information of the protocols, as per IETF standard RFC: 8285 titled “A General Mechanism for RTP Header Extensions”.
- RTP header extensions produced at the source may be ignored by the destination endpoints that do not have the knowledge to interpret and process the RTP header extensions transmitted by the source endpoint.
- Video coding domain and metadata support will now be briefly described, beginning with modern hybrid video coding.
- the interactivity and immersiveness of modern and future multimedia XR applications requires guarantees in terms of meeting packet error rate (PER) and packet delay budget (PDB) for the QoE.
- PER packet error rate
- PDB packet delay budget
- the video source jitter and wireless channel stochastic characteristics of mobile communications systems make the former challenging to meet especially for high-rate specific digital video transmissions, e.g., 4K, 3D video, 2x2K eye- buffered video etc.
- the current video source information is encoded based on 2D, 2D+Depth, or alternatively, 3D representations of video content.
- the encoded elementary stream video content is generally, regardless of the source encoder, organized into two abstraction layers meant to separate the storage and video coding domains, i.e., the network transport packetization and format, and respectively, the video coding related syntax and associated semantics of a codec.
- the first determines the bitstream format, whereas the latter specifies the contents of the video coded bitstream.
- MPEG video codec families e.g., H.264, H.265, H.266
- NAL network abstraction layer
- the NAL units may enclose both video coding layer (VCL) information, i.e., video coded content (e.g., frames, slices, tiles etc.) NALUs, and respectively, non-VCL information, i.e., parameter sets, supplemental enhancement information (SEI) messages etc..
- VCL video coding layer
- SEI supplemental enhancement information
- open-source video codec alternatives e.g., VP8/VP9, or similarly, AV1
- the AV1 bitstream is comprised of Open Bitstream Units (OBUs) and each OBU may contain one or more video coded SMM920220294-GR-NP frames, video coded tiles, non-video coded padding and non-video coded metadata as OBU_METADATA.
- OBUs Open Bitstream Units
- NALUs or alternatively, OBUs provide mechanisms that can be exploited towards the transport of metadata which is not relevant to the video coded bitstream chroma and luminance representation.
- the VCL or alternatively, the video coded information, encapsulates the video coding procedures of an encoder and compresses the source encoded video information based on some entropy coding method, e.g., context-adaptive binary arithmetic encoding (CABAC), context-adaptive variable-length coding (CAVLC) etc.
- CABAC context-adaptive binary arithmetic encoding
- CAVLC context-adaptive variable-length coding
- the coding units may be subsequently split under some tree partitioning structures, or alike hierarchical structures, as described in ITU-T standard H.264 V8 (08/2021), ITU-T standard H.265 V8 (08/2021), ITU-T standard H.266 V4 (04/2022).
- tree partitioning structures may comprise binary/ternary/quaternary trees, or under some predetermined geometrically motivated 2D segmentation patterns as described by de Rivaz, p & Haughton, (2016) in the paper titled “AV1 Bitstream & Decoding Process Specification” from the Alliance for Open Media, 182, e.g., the 10-way split.
- Encoders use visual references among such coding units to encode picture content in a differential manner based on residuals.
- the residuals are determined given the prediction modes associated with the reconstruction of information.
- Two modes of prediction are universally available as intra-prediction (shortly referred to as intra as well) or inter-prediction (or inter in short form).
- the intra mode is based on deriving and predicting residuals based on other coding units’ contents within the current picture, i.e., by computing residuals of current coding units given their adjacent coding units coded content.
- the inter mode is based, on the other hand, on deriving and predicting residuals based on coding units’ contents from other pictures, i.e., by computing residuals of current coding units given their adjacent coded pictures content.
- the residuals are then further transformed for compression using some multi- dimensional (2D/3D) spatial multimodal transform, e.g., frequency-based (i.e., Discrete Cosine Transform, or alike), or wavelet-based linear transform (e.g., Walsh Hadamard SMM920220294-GR-NP Transform, or equivalently, Discrete Wavelet Transform), to extract the most prominent frequency components of the coding units’ residuals.
- some multi-dimensional (2D/3D) spatial multimodal transform e.g., frequency-based (i.e., Discrete Cosine Transform, or alike), or wavelet-based linear transform (e.g., Walsh Hadamard SMM920220294-GR-NP Transform, or equivalently, Discrete Wavelet Transform)
- the insignificant high-frequency contributions of residuals are dropped, and the floating-point transformed representation of remaining residuals is further quantized based on some parametric quantization procedure down to a selected number of bits per sample, e.g., 8/10/12 bits
- FIG. 9 illustrates a simplified block diagram 900 of a generic video codec performing both spatial and temporal (motion) compression of a video source.
- the encoder blocks are captured within the “Encoder” tagged domain 910.
- the decoder blocks are captured within the “Decoder” tagged domain 920.
- One skilled in the art may associate the generic diagram from above describing a hybrid codec with a plethora of state-of-the-art video codecs, such as, but not limited to H.264, H.265, H.266 (generically referred to as H.26x) or VP8/VP9/AV1. As such, the concepts hereby utilized shall be considered in general sense, unless otherwise specifically clarified and reduced in scope to some codec embodiment hereafter.
- the block diagram 900 shows a raw input video frame (picture) 901 being input to a picture block partitioning function block 911 of encoder 910. A subsequent functional block 912 is illustrated as ‘spatial transform’.
- a subsequent functional block 913 is illustrated as ‘quantization’.
- a subsequent functional block 914 is illustrated as ‘entropy coding’.
- This functional block 914 outputs to video coded bitstream 902 but also to motion estimation 915 of the encoder 910.
- the motion estimation 915 outputs to inter prediction block 921 of decoder 920.
- This block 921 outputs to buffer 920, which itself outputs to recovered video frame video (picture) 903.
- Inter prediction block 921 may be switched to connect with a sum junction feeding into spatial transform 912.
- quantization block 913 may output to an inverse quantization block 926 of decoder 920.
- the inverse quantization block 926 illustrated as receiving entropy decoding 927 of video coded bitstream 902.
- the inverse SMM920220294-GR-NP quantization block 926 outputs to inverse spatial transform block 925 which, via a sum junction, feeds into a loop & visual filtering block 923, itself feeding into buffer 922.
- the block diagram 900 is illustrated by way of example, to convey the various functional blocks of a modern hybrid video codec from the perspective of both encoder and decoder operations.
- the coded residual bitstream 902 is thus encapsulated into an elementary stream as NAL units, or equivalently, as OBUs ready for storage or transmission over a network.
- OBUs are the main syntax elements of a video codec, and these may encapsulate encoded video parameters (e.g., video/sequence/picture parameter set (VPS/SPS/PPS)), one or more supplemental enhancement information (SEI) messages, or alternatively OBU metadata payloads, and encoded video headers and residuals data, (e.g., slices as partitions of a picture, or equivalently, a video frame or video tile).
- SEI Supplemental Enhancement information
- OBU metadata payloads e.g., OBU metadata payloads
- encoded video headers and residuals data e.g., slices as partitions of a picture, or equivalently, a video frame or video tile.
- the encapsulation general syntax carries information described by codec specific semantics meant to determine the usage of metadata, non-video coded data and video encoded data and aid the decoding process.
- the NAL units encapsulation syntax is composed of a header portion determining the beginning of a NAL unit and the type thereof, and a raw byte payload sequence containing the NAL unit relevant information.
- the NAL unit payload may subsequently be formed of a payload syntax or a payload specific header and an associated payload specific syntax.
- NAL units A critical subset of NAL units is formed of parameter sets, e.g., VPS, SPS, PPS, SEI messages and configuration NAL units (also known generically as non-VCL NAL units), and picture slice NAL units containing video encoded data (e.g., entropy-based arithmetic encoding) as VCL information.
- VPS parameter sets
- SPS SPS
- PPS SEI messages
- configuration NAL units also known generically as non-VCL NAL units
- picture slice NAL units containing video encoded data e.g., entropy-based arithmetic encoding
- the header 1010 contains information about the type, size, and video coding attributes and parameters of the NAL unit data enclosed information.
- the NAL unit data may be non-VCL NAL comprising video/sequence/picture parameters payload 1021, and supplemental enhancement information payload 1022, or VCL NAL comprising a frame/picture/slice payload 1023 having header 1023a and video coded payload 1023b.
- the non-VCL NALUs may include in a supplemental enhancement SMM920220294-GR-NP information payload 1022 one or more SEI messages 1022a.
- the NAL header 1010 is illustrated as comprising a NAL unit type, NAL unit byte length, video coding layer ID and temporal video coding layer ID.
- a decoder implementation may implement a bitstream parser extracting the necessary metadata information and VCL associated metadata from the NAL unit sequence 1000; decode the VCL residual coded data sequence to its transformed and quantized values; apply the inverse linear transform and recover the residual significant content; perform intra or inter prediction to reconstruct each coding unit luminance and chromatic representation; apply additional filtering and error concealment procedures; reproduce the raw picture sequence representation as video playback.
- bitstream parser extracting the necessary metadata information and VCL associated metadata from the NAL unit sequence 1000; decode the VCL residual coded data sequence to its transformed and quantized values; apply the inverse linear transform and recover the residual significant content; perform intra or inter prediction to reconstruct each coding unit luminance and chromatic representation; apply additional filtering and error concealment procedures; reproduce the raw picture sequence representation as video playback.
- These operations and procedures may happen successively, as listed, or in parallel depending on a decoder specific implementation.
- One skilled in the art should recognize that similar high-level operations are applicable to other family of video codecs
- Modern video codecs e.g., H.264, H.265, AV1, or alternatively, H.266, provide byte-aligned transport mechanisms for metadata within the video coded elementary stream, or alternatively, bitstream.
- Such non-video coded data is alternatively referred to in a video coding context as metadata since the comprised information is not related to alter the luma or chroma of the decoded frame, or alternatively, picture.
- This metadata is encapsulated as well into NALUs as SEI messages for H.26x MPEG family of codecs and in OBUs as OBU metadata for AV1.
- the metadata may have different types associated with different syntax and semantics specified additionally in the codec specifications, e.g., (ITU-T standard H.264 (08/2021), ITU-T standard H.265 V8 (08/2021), ITU-T standard H.266 V4 (04/2022), Rivaz, p & Haughton, (2016) in the paper titled “AV1 Bitstream & Decoding Process Specification” from the Alliance for Open Media, 182), or alternatively, in ITU specifications of metadata types such as for instance in ITU-T Series H specification V8 (08/2020) titled “Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services – coding of moving video: versatile supplemental enhancement information messages for coded video bitstreams”, or ITU-T Recommendation T.35, titled “Terminal provider codes notification form: available SMM920220294-GR-NP information regarding the identification of national authorities for the assignment of ITU-T recommendation T.35 terminal provider codes”.
- the user data SEI message or alternatively, OBU metadata user private data format is specified.
- all the herein discussed video codecs and their associated encoders and decoders allow the exposure via dedicated interfaces of the user data metadata type to the application layer by passthrough of either NALUs SEI messages, or alternatively, OBUs carrying OBU metadata to the application layer.
- the supported and mandatory video media codecs are AVC/H.264 with its constrained baseline profile and VP8, with additional optional support for VP9.
- Split rendering allows for the enhancement of the user experience by providing access to advanced and sophisticated rendering which would be otherwise not possible, or alternatively, require very high energy consumption on the AR/VR glasses, or alternatively, a 5G tethered UE.
- EAS Edge Application Server
- SRS Split Rendering Server
- EDN Edge Data Network
- split rendering operations may be wide, ranging from full pre- SMM920220294-GR-NP rendering on the edge to offloading partial, processing-extensive rendering operations to the edge.
- This variety is ranging from frames corresponding to 2D with single eye buffer rendering, 2D with dual eye buffers rendering, 2D with depth information rendering, and 3D scene rendering.
- the edge application server would produce a single 2D video rendering of the visual scene.
- a 2D rendering with two eye buffers and an appropriate projection may be needed.
- Other supporting streams encoding additional information, such as depth or transparency may be added too to the 2D rendering.
- the split-rendering UL traffic pattern can be summarized as: the UE, or alternatively, the AR glasses stream the pose predictions to the SRS at the EDN.
- This traffic may additionally contain an associated AR UL video stream depending on the type of applications (e.g., multi-party AR conferencing, interactive and immersive classroom AR etc.)
- the split-rendering DL traffic pattern can be summarized as: the UE, or alternatively, the AR glasses receive in return the rendered video media for the display.
- the XR runtime benefits from the rendered media to be passed together with the associated pose used for rendering to perform proper scene composition and display. For instance, the XR runtime may need to perform pose correction based on late stage reprojection, or alternatively, asynchronous time warping to match the current display time.
- XR runtime e.g., XR space handles or references and XR timestamps associated with the SRS rendering operations, for spatial and temporal synchronization.
- XR space handles or references and XR timestamps associated with the SRS rendering operations, for spatial and temporal synchronization.
- OBU metadata e.g., AV1
- SMM920220294-GR-NP immersion metadata associated with an XR application in either UL or DL.
- the format used for the payload of such user data is generic and may comprise of at least two fields.
- the first field is an identifier determining the syntax and semantics of the user data format
- the second field contains the user data information as a payload encoded according to the format identified by the first field.
- the first field is in some embodiments a UUID which needs to be signaled by a source sender to a corresponding remote receiver to allow the receiver to parse and process the user data information.
- This disclosure particularly specifies three signaling mechanisms which enable the usage of such user data over SEI messages, or alternatively, OBU metadata for interactive and immersive applications backed by split-rendering at the edge.
- the proposed signaling mechanism for the user data identifier are: application-based signaling over application plane; network-supported signaling via network application function over control plane; and SDP attribute based in-band signaling over user plane.
- the user data identifier e.g., UUID
- the disclosure herein provides an apparatus for wireless communication, comprising a processor; and a memory coupled with the processor, the processor configured to cause the apparatus to: determine a media configuration for split-rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating an encoded video stream of the video media flow comprises, as non-video coded metadata, one or more data units of multimedia immersion and interaction data; signal the media configuration to a second apparatus; establish, with the second apparatus, based at least in part on the media configuration, a multimedia split- rendering content delivery session comprising the video media flow; and use, for split- rendering the video media flow, the one or more data units of multimedia immersion and interaction data.
- the non-video coded metadata comprises: a first field comprising an identifier for a syntax and semantics representation format of the one or more data units of multimedia immersion and interaction data; and a second field comprising the one or more data units of multimedia immersion and interaction data, SMM920220294-GR-NP encoded according to the syntax and semantics representation format corresponding to the identifier of the first field.
- the identifier is a universally unique identifier ‘UUID’ indication.
- the UUID is unique to a specific application or session, or the UUID is globally unique.
- the UUID may be compliant with the ISO-IEC-11578 ANNEX A format and ISO-IEC 9834-8 Version 4 UUID, i.e., randomly generated UUID, or alike.
- the term ‘session’ comprises of a temporary and interactive, i.e., updatable, set of configurations and rules, e.g., media formats and codecs, network configuration, determining the exchange of information, including media content, between two or more endpoints connected over a network.
- the processor is configured to cause the apparatus to determine the media configuration by causing the apparatus to: determine the media configuration for a plurality of video media flows having respective encoded video streams and respective non-video coded metadata, wherein the one or more parameters of the media configuration map the identifiers of the first fields of the respective non- video coded metadata, to their respective video media flows.
- the one or more parameters of the media configuration map the first fields of the respective non-video coded metadata, to their respective video media flows, based on at least one of: an application specific stream mapping; a 5-tuple indication; and a media flow description attribute.
- a 5-tuple may describe an IP flow comprising of IP source address, IP destination address, source port, destination port, protocol identifier.
- a media description under SDP may be a media flow association to one or more UUIDs as provided in the SDP defined media description attribute.
- the processor is configured to cause the apparatus to signal the media configuration over at least one of: an application interface (which may preferably be between an application service provider ‘ASP’ of the apparatus, and a split- rendering aware application of the second apparatus); a control plane interface (preferably between a real-time communications application function ‘RTC AF’ of the apparatus, and a media session handler ‘MSH’ of the second apparatus); and a user plane interface (preferably between a split-rendering server ‘SRS’ of the apparatus, and a split- rendering client ‘SRC’ of the second apparatus).
- an application interface which may preferably be between an application service provider ‘ASP’ of the apparatus, and a split- rendering aware application of the second apparatus
- a control plane interface preferably between a real-time communications application function ‘RTC AF’ of the apparatus, and a media session handler ‘MSH’ of the second apparatus
- a user plane interface preferably between a split-rendering server ‘SRS’ of the apparatus, and a split- rendering client ‘SRC’ of the second apparatus
- the control plane interface may be a SMM920220294-GR-NP RTC-5 reference interface in 5GS; user plane interface may be SR-4, RTC-4, or may use SR-4m media centric interfacing.
- the encoded video stream is encoded/decoded using at least one video codec selected from the list of video codecs consisting of: H.264 video codec specification; H.265 video codec specification; H.266 video codec specification; and AV1 video codec specification.
- Other video codecs specification relying in part of said specification may also apply, e.g., OMAF, V-PCC or alike.
- OMAF encodes omnidirectional video content and it comprises at least one video stream encoded with AVC and HEVC
- V-PCC encodes 2D projected 3D video and it comprises a video stream encoding the projected 2D flat frames by means of AVC and HEVC.
- the video codec comprises the H.264, H.265, or H.266 codecs
- the non-video coded metadata is encapsulated as a payload of user data in one or more supplemental enhancement information ‘SEI’ messages of the type ‘user data unregistered’
- the video codec comprises the AV1 codec
- the non-video coded metadata is encapsulated as a payload of user data in one or more metadata open bitstream units ‘OBUs’ of type ‘unregistered user private data’.
- SEI Supplemental Enhancement Information
- the payload type may equal 5
- the SEI message may be prefixed or suffixed to a NALU.
- the OBU metadata_type may be ‘X’, where X can be any of 6-31.
- the one or more data units of multimedia immersion and interaction data comprise immersion and interaction data selected from the list consisting of: a user viewpoint data; a user field of view data; a user pose/orientation data; a user gesture tracking data; a user body tracking data; a user facial feature tracking data; a user action and/or user input data; a split rendering pose and spatial information; and an augmented reality object representation, comprising of at least one of a graphical description of an object and an object positional anchor.
- the one or more data units of multimedia immersion and interaction data comprises extended reality ‘XR’ multimedia immersion and interaction data.
- the one or more data units of multimedia immersion and interaction data are from one or more media sources selected from the list of media sources consisting of: one or more physical dedicated controllers; one or more red green blue ‘RGB’ cameras; one or more RGB-depth ‘RGBD’ cameras; one or more infrared ‘IR’ cameras; one or more microphones; and one or more haptic transducers.
- the processor is configured to cause the apparatus to use, for split-rendering the video media flow, the one or more data units of multimedia immersion and interaction data, for the purposes of: a late stage reprojection at a split- rendering client ‘SRC’ for displaying one or more rendered frames matching, a latest user pose or orientation information; a partial or full pre-rendering of one or more frames at a split-rendering server ‘SRS’; and/or an estimate, at a SRS, of a most probable user pose and orientation, associated with an expected display time for the one or more frames being partially or fully pre-rendered.
- SRC split- rendering client
- SRS split-rendering server
- the processor is configured to cause the apparatus to signal the media configuration upon initiation of the multimedia split-rendering content delivery session and/or when updating the multimedia split-rendering content delivery session.
- the processor is further configured to cause the apparatus to determine the media configuration based on at least one of: an application configuration of an application hosted on the second apparatus, wherein the application configuration comprises one or more media capabilities of the second apparatus; and/or an application service provider ‘ASP’ configuration, provisioned to at least one of: a real- time communications application function ‘RTC AF’; a provisioning function; a split- rendering application function ‘SR AF’; and/or a configuration function;
- the application configuration may be preconfigured with supported UUIDs; and/or transport may be automatically enabled if a UUID is received.
- FIG. 11 illustrates an embodiment 1100 of a method of wireless communication in a wireless communication system.
- a first step 1110 comprises, determining a media configuration for split- rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating an encoded video stream of the video media flow comprises, as non-video coded metadata, one or more data units of multimedia immersion and interaction data.
- a further step 1120 comprises, signaling the media configuration to a second apparatus.
- a further step 1130 comprises establishing, with the second apparatus, based at least in part on the media configuration, a multimedia split-rendering content delivery session comprising the video media flow.
- a further step 1140 comprises using, for split-rendering the video media flow, the one or more data units of multimedia immersion and interaction data.
- the method 1100 may be performed by a processor executing program code, for example, a microcontroller, a microprocessor, a CPU, a GPU, an auxiliary processing unit, a FPGA, or the like.
- the non-video coded metadata comprises: a first field comprising an identifier for a syntax and semantics representation format of the one or more data units of multimedia immersion and interaction data; and a second field comprising the one or more data units of multimedia immersion and interaction data, encoded according to the syntax and semantics representation format corresponding to the identifier of the first field.
- the identifier is a universally unique identifier ‘UUID’ indication.
- the UUID is unique to a specific application or session, or the UUID is globally unique.
- the UUID may be compliant with ISO-IEC-11578 ANNEX A format and ISO-IEC 9834-8 Version 4 UUID, i.e., randomly generated UUID, or alike.
- a ‘session’ comprises of a temporary and interactive, i.e., updatable, set of configurations and rules, e.g., media formats and codecs, network configuration, determining the exchange of information, including media content, between two or more endpoints connected over a network.
- the determining the media configuration comprises: determining the media configuration for a plurality of video media flows comprising respective encoded video streams and respective non-video coded metadata, wherein the one or more parameters of the media configuration map the identifiers of the first fields of the respective non-video coded metadata, to their respective video media flows.
- the one or more parameters of the media configuration map the first fields of the respective non-video coded metadata, to their respective video media flows, based on at least one of: an application specific stream mapping; a 5-tuple indication; and a media flow description attribute.
- a 5-tuple may describe an IP flow comprising of IP source address, IP destination address, source port, destination port, protocol identifier, for instance.
- a media description under SDP may comprise a media SMM920220294-GR-NP flow association to one or more UUIDs as provided in the SDP media description attribute etc.
- the signaling the media configuration comprises signaling over at least one of: an application interface (preferably between an application service provider ‘ASP’ of the apparatus, and a split-rendering aware application of the second apparatus); a control plane interface (preferably between a real-time communications application function ‘RTC AF’ of the apparatus, and a media session handler ‘MSH’ of the second apparatus); and a user plane interface (preferably between a split-rendering server ‘SRS’ of the apparatus, and a split-rendering client ‘SRC’ of the second apparatus).
- an application interface preferably between an application service provider ‘ASP’ of the apparatus, and a split-rendering aware application of the second apparatus
- a control plane interface preferably between a real-time communications application function ‘RTC AF’ of the apparatus, and a media session handler ‘MSH’ of the second apparatus
- the control plane interface may be RTC-5 reference interface in 5GS; user plane interface may be SR-4, RTC-4, or may use SR-4m media centric interfacing.
- the encoded video stream is encoded/decoded using at least one video codec selected from the list of video codecs consisting of: H.264 video codec specification; H.265 video codec specification; H.266 video codec specification; and AV1 video codec specification.
- Other video codecs specification relying in part of said specification may also apply, e.g., OMAF, V-PCC or alike.
- OMAF encodes omnidirectional video content and it comprises at least one video stream encoded with AVC and HEVC
- V-PCC encodes 2D projected 3D video and it comprises a video stream encoding the projected 2D flat frames by means of AVC and HEVC.
- the video codec comprises the H.264, H.265, or H.266 codecs
- the non-video coded metadata is encapsulated as a payload of user data in one or more supplemental enhancement information ‘SEI’ messages of the type ‘user data unregistered’
- the video codec comprises the AV1 codec
- the non-video coded metadata is encapsulated as a payload of user data in one or more metadata open bitstream units ‘OBUs’ of type ‘unregistered user private data’.
- the payload type may be 5, furthermore the SEI message may be prefixed or suffixed to a NALU.
- the OBU metadata_type may be ’X’, where X can be any of 6-31.
- the one or more data units of multimedia immersion and interaction data comprise immersion and interaction data selected from the list consisting of: a user viewpoint data; a user field of view data; a user pose/orientation data; a user gesture tracking data; a user body tracking data; a user facial feature tracking data; a user action and/or user input data; a split rendering pose and spatial information; and an augmented reality object representation, comprising of at least one of a graphical description of an object and an object positional anchor.
- the one or more data units of multimedia immersion and interaction data comprise extended reality ‘XR’ multimedia immersion and interaction data.
- the one or more data units of multimedia immersion and interaction data are from one or more media sources selected from the list of media sources consisting of: one or more physical dedicated controllers; one or more red green blue ‘RGB’ cameras; one or more RGB-depth ‘RGBD’ cameras; one or more infrared ‘IR’ cameras; one or more microphones; and one or more haptic transducers.
- Some embodiments comprise using, for split-rendering the video media flow, the one or more data units of multimedia immersion and interaction data, for the purposes of: a late stage reprojection at a split-rendering client ‘SRC’ for displaying one or more rendered frames matching, a latest user pose or orientation information; a partial or full pre-rendering of one or more frames at a split-rendering server ‘SRS’; and/or an estimate, at a SRS, of a most probable user pose and orientation, associated with an expected display time for the one or more frames being partially or fully pre-rendered.
- SRC split stage reprojection at a split-rendering client ‘SRC’ for displaying one or more rendered frames matching, a latest user pose or orientation information
- SRS split-rendering server
- an estimate, at a SRS of a most probable user pose and orientation, associated with an expected display time for the one or more frames being partially or fully pre-rendered.
- the determining the media configuration comprises determining the media configuration based on at least one of: an application configuration of an application hosted on the second apparatus, wherein the application configuration comprises one or more media capabilities of second apparatus; and/or an application service provider ‘ASP’ configuration, provisioned to at least one of a real- time communications application function ‘RTC AF’; a provisioning function; a split- rendering application function ‘SR AF’; and/or a configuration function;
- the application configuration may be preconfigured with supported UUIDs and/or transport may be automatically enabled if a UUID is received.
- the method is performed by a network entity/node, and the second apparatus is a UE.
- the disclosure herein further provides an apparatus for wireless communication in a wireless communication system, comprising a processor; and a memory coupled with the processor, the processor configured to cause the apparatus to: receive, from a first apparatus, a media configuration for split-rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating an encoded video SMM920220294-GR-NP stream of the video media flow comprises, as non-video coded metadata, one or more data units of multimedia immersion and interaction data; configure the apparatus, using the media configuration, to receive the video media flow; decode the encoded video stream of the video media flow, wherein the decoding comprises extracting the non- video coded metadata; and consume, from the non-video coded metadata, the one or more data units of multimedia immersion and interaction data.
- the non-video coded metadata comprises: a first field comprising an identifier for a syntax and semantics representation format of the one or more data units of multimedia immersion and interaction data; and a second field comprising the one or more data units of multimedia immersion and interaction data, encoded according to the syntax and semantics representation format corresponding to the identifier of the first field.
- the identifier is a universally unique identifier ‘UUID’ indication.
- UUID is unique to a specific application or session, or the UUID is globally unique.
- the UUID may be compliant with ISO-IEC-11578 ANNEX A format and ISO-IEC 9834-8 Version 4 UUID, i.e., randomly generated UUID, or alike.
- a ‘session’ comprises of a temporary and interactive, i.e., updatable, set of configurations and rules, e.g., media formats and codecs, network configuration, determining the exchange of information, including media content, between two or more endpoints connected over a network.
- the media configuration is for a plurality of video media flows comprising respective encoded video streams and respective non-video coded metadata, wherein the one or more parameters of the media configuration map the identifiers of the first fields of the respective non-video coded metadata, to their respective video media flows.
- the one or more parameters of the media configuration map the first fields of the respective non-video coded metadata, to their respective video media flows, based on at least one of: an application specific stream mapping; a 5-tuple indication; and a media flow description attribute.
- a 5-tuple may describe an IP flow comprising of IP source address, IP destination address, source port, destination port, protocol identifier.
- a media description under SDP may be a media flow association to one or more UUIDs as provided in the SDP media description attribute etc.
- SMM920220294-GR-NP the processor is configured to cause the apparatus to receive the media configuration over at least one of: an application interface (preferably between an application service provider ‘ASP’ of the apparatus, and a split-rendering aware application of the second apparatus); a control plane interface (preferably between a real-time communications application function ‘RTC AF’ of the apparatus, and a media session handler ‘MSH’ of the second apparatus); and a user plane interface (preferably between a split-rendering server ‘SRS’ of the apparatus, and a split-rendering client ‘SRC’ of the second apparatus).
- an application interface preferably between an application service provider ‘ASP’ of the apparatus, and a split-rendering aware application of the second apparatus
- RTC AF real-time communications application function
- MSH media session handler
- SRC split-rendering client
- the control plane interface may be a RTC-5 reference interface in 5GS.
- the user plane interface may be SR-4, RTC-4, or may use SR-4m media centric interfacing.
- the encoded video stream is encoded/decoded using at least one video codec selected from the list of video codecs consisting of: H.264 video codec specification; H.265 video codec specification; H.266 video codec specification; and AV1 video codec specification.
- Other video codecs specification relying in part of said specification may also apply, e.g., OMAF, V-PCC or alike.
- OMAF encodes omnidirectional video content and it comprises at least one video stream encoded with AVC and HEVC
- V-PCC encodes 2D projected 3D video and it comprises a video stream encoding the projected 2D flat frames by means of AVC and HEVC.
- the video codec comprises the H.264, H.265, or H.266 codecs
- the non-video coded metadata is encapsulated as a payload of user data in one or more supplemental enhancement information ‘SEI’ messages of the type ‘user data unregistered’
- the video codec comprises the AV1 codec
- the non-video coded metadata is encapsulated as a payload of user data in one or more metadata open bitstream units ‘OBUs’ of type ‘unregistered user private data’.
- the payload type may be 5, furthermore the SEI message may be prefixed or suffixed to a NALU.
- the OBU metadata_type may be X, where X can be any of 6-31.
- the one or more data units of multimedia immersion and interaction data comprise immersion and interaction data selected from the list consisting of: a user viewpoint data; a user field of view data; a user pose/orientation data; a user gesture tracking data; a user body tracking data; a user facial feature tracking data; a user action and/or user input data; a split rendering pose and spatial information; and an augmented reality object representation, comprising of at least one of a graphical description of an object and an object positional anchor.
- the one or more data units of multimedia immersion and interaction data comprise extended reality ‘XR’ multimedia immersion and interaction data.
- the one or more data units of multimedia immersion and interaction data are from one or more media sources selected from the list of media sources consisting of: one or more physical dedicated controllers; one or more red green blue ‘RGB’ cameras; one or more RGB-depth ‘RGBD’ cameras; one or more infrared ‘IR’ cameras; one or more microphones; and one or more haptic transducers.
- the processor is configured to cause the apparatus to use, for split-rendering the video media flow, the one or more data units of multimedia immersion and interaction data, for the purposes of: a late stage reprojection at a split- rendering client ‘SRC’ for displaying one or more rendered frames matching, a latest user pose or orientation information.
- the processor is configured to cause the apparatus to receive the media configuration upon initiation of the multimedia split-rendering content delivery session and/or when updating the multimedia split-rendering content delivery session.
- the media configuration is based on at least one of: an application configuration of an application hosted on the apparatus, wherein the application configuration comprises one or more media capabilities of the apparatus; and/or an application service provider ‘ASP’ configuration, provisioned to at least one of: a real-time communications application function ‘RTC AF’; a provisioning function; [0176] a split-rendering application function ‘SR AF’; and/or a configuration function.
- the application configuration may be preconfigured with supported UUIDs and/or transport may be automatically enabled if a UUID is received.
- the first apparatus is a network entity/node, whereas the apparatus is a UE.
- FIG. 12 illustrates an embodiment of a method 1200 for wireless communication.
- a first step 1210 comprises receiving, from a first apparatus, a media configuration for split-rendering of a video media flow, wherein the media configuration comprises one or more parameters indicating an encoded video stream of the video media flow comprises, as non-video coded metadata, one or more data units of multimedia immersion and interaction data.
- SMM920220294-GR-NP [0180]
- a further step 1220 comprises configuring the apparatus, using the media configuration, to receive the video media flow.
- a further step 1230 comprises decoding the encoded video stream of the video media flow, wherein the decoding comprises extracting the non-video coded metadata.
- a further step 1240 comprises consuming, from the non-video coded metadata, the one or more data units of multimedia immersion and interaction data.
- the method 1200 may be performed by a processor executing program code, for example, a microcontroller, a microprocessor, a CPU, a GPU, an auxiliary processing unit, a FPGA, or the like.
- the non-video coded metadata comprises: a first field comprising an identifier for a syntax and semantics representation format of the one or more data units of multimedia immersion and interaction data; and a second field comprising the one or more data units of multimedia immersion and interaction data, encoded according to the syntax and semantics representation format corresponding to the identifier of the first field.
- the identifier is a universally unique identifier ‘UUID’ indication.
- the UUID is unique to a specific application or session, or the UUID is globally unique.
- the UUID may be compliant with the ISO-IEC-11578 ANNEX A format and ISO-IEC 9834-8 Version 4 UUID, i.e., randomly generated UUID, or alike.
- a ‘session’ comprises of a temporary and interactive, i.e., updatable, set of configurations and rules, e.g., media formats and codecs, network configuration, determining the exchange of information, including media content, between two or more endpoints connected over a network.
- the media configuration is for a plurality of video media flows comprising respective encoded video streams and respective non-video coded metadata, wherein the one or more parameters of the media configuration map the identifiers of the first fields of the respective non-video coded metadata, to their respective video media flows.
- the one or more parameters of the media configuration map the first fields of the respective non-video coded metadata, to their respective video media flows, based on at least one of: an application specific stream mapping; a 5-tuple indication; and a media flow description attribute.
- SMM920220294-GR-NP [0190]
- a 5-tuple may describe an IP flow comprising of IP source address, IP destination address, source port, destination port, protocol identifier.
- a media description under SDP may comprise a media flow association to a description attribute as provided in the SDP media description attribute etc.
- the receiving the media configuration comprises receiving over at least one of: an application interface (preferably between an application service provider ‘ASP’ of the apparatus, and a split-rendering aware application of the second apparatus); a control plane interface (preferably between a real-time communications application function ‘RTC AF’ of the apparatus, and a media session handler ‘MSH’ of the second apparatus); and a user plane interface (preferably between a split-rendering server ‘SRS’ of the apparatus, and a split-rendering client ‘SRC’ of the second apparatus).
- the control plane interface may be RTC-5 reference interface in 5GS.
- the user plane interface may be SR-4, RTC-4, or may use SR-4m media centric interfacing.
- the decoding uses at least one video codec selected from the list of video codecs consisting of: H.264 video codec specification; H.265 video codec specification; H.266 video codec specification; and AV1 video codec specification.
- Some other video codecs’ specification relying in part of said specification may also apply, e.g., OMAF, V-PCC or alike.
- OMAF encodes omnidirectional video content and it comprises at least one video stream encoded with AVC and HEVC
- V-PCC encodes 2D projected 3D video and it comprises a video stream encoding the projected 2D flat frames by means of AVC and HEVC.
- the video codec comprises the H.264, H.265, or H.266 codecs
- the non-video coded metadata is encapsulated as a payload of user data in one or more supplemental enhancement information ‘SEI’ messages of the type ‘user data unregistered’
- the video codec comprises the AV1 codec
- the non-video coded metadata is encapsulated as a payload of user data in one or more metadata open bitstream units ‘OBUs’ of type ‘unregistered user private data’.
- SEI messaging the payload type may be 5, furthermore the SEI message may be prefixed or suffixed to a NALU.
- the OBU metadata_type may be ‘X’, where X can be any of 6-31.
- the one or more data units of multimedia immersion and interaction data comprise immersion and interaction data selected from the list consisting of: a user viewpoint data; a user field of view data; a user pose/orientation data; a user SMM920220294-GR-NP gesture tracking data; a user body tracking data; a user facial feature tracking data; a user action and/or user input data; a split rendering pose and spatial information; and an augmented reality object representation, comprising of at least one of a graphical description of an object and an object positional anchor.
- the one or more data units of multimedia immersion and interaction data comprise extended reality ‘XR’ multimedia immersion and interaction data.
- the one or more data units of multimedia immersion and interaction data are from one or more media sources selected from the list of media sources consisting of: one or more physical dedicated controllers; one or more red green blue ‘RGB’ cameras; one or more RGB-depth ‘RGBD’ cameras; one or more infrared ‘IR’ cameras; one or more microphones; and one or more haptic transducers.
- Some embodiments comprise using, for split-rendering the video media flow, the one or more data units of multimedia immersion and interaction data, for the purposes of: a late stage reprojection at a split-rendering client ‘SRC’ for displaying one or more rendered frames matching, a latest user pose or orientation information.
- Some embodiments comprising receiving the media configuration upon initiation of the multimedia split-rendering content delivery session and/or when updating the multimedia split-rendering content delivery session.
- the media configuration is based on at least one of: an application configuration of an application hosted on the apparatus, wherein the application configuration comprises one or more media capabilities of the apparatus; and/or an application service provider ‘ASP’ configuration, provisioned to at least one of a real-time communications application function ‘RTC AF’; a provisioning function; a split-rendering application function ‘SR AF’; and/or a configuration function.
- the application configuration may be preconfigured with supported UUIDs and/or transport may be automatically enabled if a UUID is received.
- the first apparatus is a network entity/node, whereas the apparatus is a UE.
- FIG. 13a provides an illustration of an embodiment 1300 of a split-rendering architecture for interactive and immersive multimedia applications, that is supported by split-rendering at the edge.
- the embodiment 1300 shows a UE 1310, an EN 1320 and a DN 1330 interfacing over various interfaces.
- the UE 1310 is illustrated as comprising a split-rendering-aware application client 1311,an MSH 1312 which may comprise an EEC 1312a, a split-rendering client (SRC) 1313, and an XR runtime 1314.
- the split-rendering-aware application client 1311 interfaces with XR runtime 1314.
- the split-rendering-aware application client 1311 interfaces with SRC 1313 over, for instance, SR-7.
- the split-rendering-aware application client 1311 interfaces with MSH 1312 over, for instance, SR-6, and may interface with EEC 1312a over EDGE-5.
- the MSH 1312 interfaces with SRC 1313 over, for instance, SR-7 and SR-6.
- the SRC 1313 interfaces with XR runtime 1314.
- the EN 1320 is illustrated as comprising a SR AF 1321 comprising a configuration function 1321a, a provisioning function 1321b, and an EES 1321c.
- the EES may be distributed over multiple instances 1321c which interface with each other over EDGE-9, for instance.
- the EN 1320 is further illustrated as comprising an RTC AS 1322, comprising a signaling server 1322a, and an SRS 1322b which may comprise EAS.
- the SR AF 1321 interfaces with RTC AS 1322 over, for instance, SR-3.
- the EES 1321c of SR AF 1321 interfaces with SRS 1322b comprising EAS over EDGE-3.
- the DN 1330 is illustrated as comprising an ASP 1331.
- the ASP 1331 interfaces with provisioning function 1321b of SR AF 1321 over SR-1.
- the ASP 1331 interfaces with SRS 1322b comprising EAS over, for instance, SR-2.
- the split-rendering-aware application client 1311 of UE 1310 interfaces with ASP 1331 of DN 1330 over, for instance, SR-8.
- the MSH 1312 of UE 1310 interfaces with SR AF 1321 of EN 1320 over, for instance, SR-5.
- the EEC 1312a of MSH 1312 may further interface with EES 1321c of SR AF 1321 over EDGE 1 and/or EDGE 4 (for instance via an ECS).
- the SRC 1313 of UE 1310 interfaces with the SRS 1322b and signaling server 1322a of RTC AS 1322 in EN 1320, over, for instance, SR-4 (comprised of SR-4m, SR-4s, for instance).
- SR-4 compact of SR-4m, SR-4s, for instance.
- FIG. 13b provides an illustration of the high-level split-rendering flows in an embodiment. Illustrated in the figure are the split-rendering-aware application 1311, MSH 1312, SRC 1313 (illustrated as comprising a scene manager, XR source manager and media access function), SRS 1322b, SR AF 1321, application service provider 1331, and XR runtime 1314.
- SR AF split-rendering application function
- RTC AF generic real-time communications application function
- the ASP 1331 provisions the SR AF 1321 via a provisioning request for a split-rendering management session.
- the provisioning may be performed in some embodiments over the SR-1 reference point, or alternatively, SR-1 APIs.
- the SR AF- 1321 determines a split-rendering configuration based in part on the information received from the ASP 1331.
- an edge enabled SR AF 1321 may support additionally Edge Enabler Server (EES).
- EES Edge Enabler Server
- the EES logic may be distributed over multiple edge DNs, whereby the context and coordination of the distributed EES instances is coordinated over the EDGE-9 interface.
- the EES may further communicate with and register EAS split-rendering functionality.
- the split-rendering server 1322b may represent an instantiation of the EAS for split-rendering and may communicate with the SR AF 1321 over the EDGE-3 interface. This may imply for example registration/deregistration of EAS instances, access network capabilities exposure, QoS management notifications and reports (e.g., bitrate adaptation notifications etc.). This step 1301 is illustrated as “split rendering provisioning”.
- the split-rendering-aware application 1311 acquires the Service Access Information (SAI) from the ASP 1331.
- SAI is a set of parameters and addresses that are needed by a client to activate the reception of one or more DL/UL media sessions, perform dynamic policy invocation, consumption/metrics reporting, and request SR AF 1321 assistance.
- the acquisition of SAI may be performed in some applications by means of control plane signaling over SR-5 (or alike RTC-5) from the SR AF 1321 to the MSH 1312 followed by exposure of the acquired SAI by the MSH 1312 SMM920220294-GR-NP to the application 1311 over SR-6 (or alike RTC-6) interface.
- the SR-5 interface may in some embodiments comprise EDGE-1 functionality aiding an Edge Enabler Client instantiated by the MSH 1312 with the registration to the EES at the SR AF 1321, and respectively, EDGE-4 functionality enabling the EEC to access configuration of edge resources (e.g., EAS configuration information) from an Edge Configuration Server (ECS).
- edge resources e.g., EAS configuration information
- ECS Edge Configuration Server
- the SAI may be exposed directly by the ASP 1331 (as acquired after the SR-1 provisioning invocation) over the application specific SR-8 (or similar RTC-8) interface. This step 1302 is illustrated as “service access information (SAI) acquisition”.
- SAI service access information
- step 1303a proceeds with the application 1311 requesting the split-rendering assistance to the UE SRC 1313 over the SR-7 API.
- the SRC 1313 communicates with the MSH 1312 over the SR-6 API to determine the client media capabilities and functions (e.g., XR runtime API, XR runtime rendering capabilities etc.), as well as available SRS instances as per the SAI configuration.
- the SRC 1313 then negotiates the media-centric split configuration with the SRS 1322b at the edge.
- This may imply in some embodiments the split management negotiation via SDP procedures or dedicated split-rendering management signaling (e.g., as add-on to WebRTC specific signaling etc.).
- SDP dedicated split-rendering management signaling
- the split is negotiated between the SRC 1313 and SRS 1322b based on the SDP protocol, whereby the SDP offer/answer procedures are used to determine the media streams, media formats and multiplexing supported by the SRC 1313 and SRS 1322b during the split-rendering session.
- the SRS 1322b acknowledges to the SRC 1313 split configuration over the user plane interface SR-4 (or alike RTC-4), and the split-rendering media delivery session is ready to start. As a result, the SRC 1313 completes the application request for split- rendering over the SR-7 interface.
- the step 1303a is illustrated as “client-driven split management”.
- the step 1303b proceeds with the SRS 1322b requesting the SRC 1313 for a split- rendering session over the SR-4 (or alike RTC-4) interface.
- this SMM920220294-GR-NP may be in part achieved over a signaling server functionality (i.e., implemented by a dedicated signaling server or alike), whereas in other embodiments this may be left to an SDP media session negotiation based on offer/answer procedure.
- the SRC 1313 queries the MSH 1312 about the UE media capabilities given the provisioned split-rendering configuration and available SAI.
- the MSH 1312 replies to the SRC 1313 over the SR-6 with the available media capabilities and information.
- the SRC 1313 replies to the SRS 1322b request over the SR-4 (or alike RTC-4) interface, and negotiates the media-centric split management.
- a further step 1304, comprises the media delivery. In some embodiments this includes the establishment of a media delivery split-rendering session by the SRC 1313.
- the SRC 1313 may request the SRS 1322b to create a split rendering session given the determined split-rendering configuration profile from the previous step.
- the SRC 1313 is formed of at least a scene manager shim functionality, an XR source manager and a media access function.
- Such an SRC 1313 may thus trigger by means of the scene manager functionality, or alternatively, by means of the MAF functionality the request to the SRS 1322b to create the split rendering session.
- the SRS 1322b responds to acknowledge the request with a split-rendering description of the split-rendering configuration profile for the session.
- the SRC 1313 establishes connection to the SRS 1322b and the split-rendering sessions is initiated.
- this may apply additional media-centric signaling given a signaling server (e.g., WebRTC over WebSockets or other signaling mechanisms) over SR-4s, as a signaling subinterface of SR-4.
- a signaling server e.g., WebRTC over WebSockets or other signaling mechanisms
- media is served as configured between the SRC 1313 and SRS 1322b over the SR-4 over both UL and DL directions.
- This step 1304 is illustrated as “media delivery setup”.
- a further step 1305 following with a rendering loop includes in some embodiments the SRC 1313 receiving, or alternatively acquiring, from the XR runtime pose information and user SMM920220294-GR-NP actions.
- the latter information is complemented by an UL video and/or audio stream (e.g., as for AR conferencing or AR immersive and interactive applications).
- the pose, user action and any media streams are transmitted in UL to the SRS 1322b over SR-4m (i.e., the user plane media-centric subinterface of SR-4).
- the SRS 1322b processes the received information, the information in its own service buffer and renders at least in part the contents of the next one or more frames for the split- rendering-aware application.
- the SRS 1322b may fuse together one or more pose and user actions to estimate the pose information as close as possible to the expected display time of the frame. In such an embodiment this estimate is used to render, or alternatively, pre-render the next frame.
- the SRS 1322b may select the most recent pose information to expected display time of the next frame for the rendering/pre-rendering operation.
- the SRS 1322b transmits the rendered frames to the SRC 1313 over the SR-4m.
- the rendered frames may include additionally interaction and immersion metadata (e.g., pose information, user actions/gestures) that were used by the SRS 1322b to render the frames. This data is sent together with the video encoded rendered frames as parts of a video elementary stream associated of the split-rendering session.
- this interaction and immersion metadata may be embedded in the video elementary stream as a user specific metadata, e.g., as SEI messages for H.264, H.265, H.266, or alternatively, as OBU metadata for AV1.
- the media streams are then transmitted by the SRS 1322b to the SRC 1313.
- the MAF functions and video codecs decode the media streams and expose the user data embedded into the elementary streams further to the XR runtime 1314, or alternatively, to the split-rendering-aware application 1311 over the SR-7 interface.
- the XR runtime 1314 displays the rendered frames and uses the embedded interaction and immersion metadata for late stage reprojection and asynchronous time- warping in correcting any errors between the current pose information to the XR runtime 1314 at the display time, and the SRS 1322b estimated/used pose and user action information during the rendering operation.
- the payload of interaction and immersion metadata being communicated may contain at least two fields.
- the first field acts as a unique type identifier, i.e., a UUID, determining the syntax and semantics of the corresponding information payload located in a second field.
- the second field carries the interaction and immersion data, data embedded as metadata into the video coded elementary stream.
- the second field SMM920220294-GR-NP syntax and semantics may be determined at least in part based on the first field.
- This may imply an indexed search (e.g., in a list of supported formats for interaction and immersion metadata by a particular communication endpoint, like an UE, or alternatively, an AS), a data repository, or alternatively, registry search (e.g., a query against an Internet based registry of one or more formats for interaction and immersion metadata), or a selection of pre-configured resources (e.g., a selection of an application determined format for interaction and immersion metadata).
- Figure 14 illustrates a representation 1400 of multimedia interaction and immersion user data as metadata within a video coded elementary stream for MPEG H- 26x family of video codecs. Illustrated for a given NAL unit 1410 is a NAL header 1411 and a NAL payload (Raw Bytes) 1412.
- the NAL payload 1412 comprises a NAL SEI Raw Byte Sequence Payload 1413 which itself comprises first and second SEI messages 1413a and 1413b.
- Figure 15 illustrates a representation 1500 of multimedia interaction and immersion user data as metadata within a video coded elementary stream for the AV1 video codec.
- the video decoder e.g., of a MAF within an SRC placed at the UE for DL split-rendering traffic, or alternatively, of an ASP placed at the SRS for UL split-rendering traffic, may passthrough the user data metadata (comprising the interaction an immersion data) embedded into the video elementary stream (e.g., H.264, H.265, H.266, AV1 or alike).
- the video decoder may expose the user data to other functional blocks (e.g., XR runtime, split-rendering-aware application, SRS).
- this may be done by the SRC over the SR-7 interface, or alternatively, API for any registered consumers, e.g., the split-rendering-aware application, or alternatively, the XR runtime. This may happen in some embodiments once at least one of the one or more UUID have been configured, and the feature of interaction and immersion metadata transport over video coded elementary streams has been enabled. [0224] To enable the filtering and output of interaction and immersion user data corresponding to metadata from video elementary streams, the signaling of UUIDs and associated metadata formats used by split-rendering and carried over the media-centric user plane is therefore necessary.
- UUIDs may generically stand as reference to an identifier pertaining to the determination of an interaction and immersion user data type and payload format. If not explicitly specified by examples, a UUID shall be treated herein as a generic identifier.
- Certain embodiments will now be described with reference to the signaling utilized. In particular, application-based; network-supported; and SDP-based signaling embodiments will be described. Application-based signaling will be described in the first instance.
- the ASP may signal to a split-rendering-aware application the transport enablement of interaction and immersion user data as metadata over one or more video elementary stream.
- the media streams are part of the user plane media- centric traffic associated with the DL split-rendering video traffic (e.g., split-rendered video streams comprising single or dual eye buffered of an XR application), or alternatively, with the UL video traffic (e.g., the UL one or more video streams associated with an AR interactive and immersive application).
- an application may activate this feature based on some application specific configuration (e.g., a JSON/YAML/XML media format description and configuration) which is shared between the server, i.e., the ASP application infrastructure, and the client, the split-rendering-aware application.
- the interface used to convey this configuration may comprise in some examples a proprietary implementation or signaling protocol served over reliable communication channel, such as an SR-8 interface as outlined in Figures 13a and 13b. Some example implementations may use WebSockets, HTTP methods, SCTP or other reliable and acknowledged messaging protocols to communicate this information.
- the configuration may further contain a list of one or more UUIDs supported by the application each corresponding to an interaction and immersion user data payload format.
- these formats may be dependent on the XR runtime used by the UE i.e., corresponding to the SRC in the split- rendering setup. For example, this would be the case between an OpenXR hand tracking extension format, and a Microsoft ® HoloLens2 specific hand tracking format.
- the formats may be commonly encoded to correspond to one or more XR runtimes, e.g., a set of OpenXR abstract formats valid for one or more devices.
- the UUIDs may be corresponding to one or more types of interaction and immersion user data payloads.
- a UUID may correspond in some examples to vanilla pose information, e.g., OpenXR’s XrPosef, corresponding to user head tracking, and another UUID may correspond to a set of multiplex poses, locations, and velocities, each corresponding to a hand joint associated with user hand tracking, e.g., OpenXR’s XR_EXT_hand_tracking.
- the application logic may be static with respect to the supported UUIDs and XR runtimes by an application.
- the split- rendering-aware application is configured with a static configuration of the supported list of UUIDs.
- the feature of transporting the interaction and immersion user data of a split-rendering application as metadata over video coded elementary streams is automatically enabled when at least one UUID is provided for filtering.
- a split-rendering-aware application enabled and configured with the interaction and immersion user data UUIDs may further share its configuration in some embodiments to the SRC.
- the SRC may apply the received configuration to filter by UUID the corresponding payloads of interaction and immersion user data post video decoding. These data payloads may further be exposed via other interfaces.
- an SRS may request from the SR AF, which may be directly indicated by the ASP, e.g., over SR-1 interface, the configuration information regarding the transport enablement of interaction and immersion user data as metadata over video elementary streams, and the corresponding UUIDs of such metadata types and payloads.
- the ASP signaling to application of the configuration information regarding the transport enablement of interaction and immersion user data as metadata over video elementary streams, and the corresponding UUIDs of such metadata types and payloads may additionally comprise the mapping of the UUIDs to video elementary streams.
- this may be performed by means of an application-specific stream mapping (e.g., based on application media flow identifiers), whereas in another example this may be performed by means of 5-tuple information (src addr, dst addr, srd port, dst port, protocol identifier) of the available media flows.
- one or more application functions may signal to the application the transport enablement of interaction and immersion user data as metadata over one or more video elementary streams.
- the signaling of this configuration may comprise additionally of UUIDs determining the type and payload format of the interaction and immersion user data to be embedded as metadata over the video elementary stream.
- such a container may be indicated by an SR AF from the split-rendering edge DN to the MSH on a UE by means of interface SR-5.
- this information is used to indicate in part the media configuration (e.g., used video codecs, interaction and immersion types and formats of the metadata embedded in the video coded elementary streams).
- the SR-5 may comprise functionality specific to UL or DL media streaming corresponding to M5d or M5u interfaces of 5GMS architecture, whereas in other examples, the SR-5 may comprise functionality specific to RTC-5 interface corresponding to the 5GS real-time communications (RTC) architecture.
- RTC real-time communications
- the MSH containing the configuration information regarding the transport enablement of interaction and immersion user data as metadata over one or more video elementary streams, and the corresponding UUIDs of such metadata may expose via an interface this configuration to the SRC.
- the SRC may request in turn this information upon starting or updating a split-rendering media delivery session to appropriately handle, and further expose the interaction and immersion user data to other functional blocks.
- this communication may happen over the SR-6 APIs exposing the MSH configuration information regarding interaction and immersion metadata transport and formats to the SRC.
- Such an interface may be implemented in some examples by typical HTTP methods, e.g., GET, or any other request-response based protocols.
- an SRS may request from the SR AF via an available API, e.g., SR-3, the configuration information regarding the transport enablement of interaction and immersion user data as metadata over video elementary streams, and the corresponding UUIDs of such metadata types and payloads.
- an available API e.g., SR-3
- the SR AF signaling to the MSH of the configuration information regarding the transport enablement of interaction and immersion user data as metadata over video elementary streams, and the corresponding UUIDs of such metadata types and payloads may additionally comprise the mapping of the UUIDs to SMM920220294-GR-NP video elementary streams.
- the split-rendering server may signal to the application a list of one or more UUIDs corresponding to the transport enablement of one or more types and formats of the interaction and immersion user data as metadata over one or more video elementary streams.
- the split-rendering server signals to this end this configuration information per each video elementary stream, thus indicating which UUIDs are supported by each video coded elementary stream.
- the split-rendering server signals the latter in-band over the media-centric user plane to its corresponding split-rendering client.
- the protocol used to signal this configuration is SDP and the signaling and media negotiation of supported UUIDs between the server and the client is based on the SDP offer/answer procedure.
- the SDP offer/answer is performed before the split-rendering content delivery media session establishment, or alternatively, update.
- the SRS provides an SDP offer to the SRC.
- the SRC parses the offer, identifies the SDP attributes associated with the interaction and immersion user data and determines based on the parsed value whether the latter are supported or not.
- the SRC copies in the SDP answer to the SRS the supported UUIDs configuration corresponding to the interaction and immersion user data types and formats. This configuration is consequently applied to the DL traffic of a split-rendering content delivery session.
- UL split-rendering traffic may comprise additionally a video stream (e.g., an AR application)
- the SDP offer/answer procedure is reciprocated.
- the SRC provides an SDP offer to the SRS with supported UUIDs for the video elementary stream embedded interaction and immersion metadata.
- the SRS processes the SDP offer and provides in turn its response in an SDP answer. If the SRS accepts the SDP UUIDs attributes offer, it copies these to the corresponding SDP answer which is sent back to the SRC.
- the split-rendering media delivery session starts with the determined SDP offer/answer negotiation outcome and the supported SMM920220294-GR-NP UUIDs types and corresponding payloads are embedded as interaction and immersion metadata to the video coded elementary stream in UL.
- the SDP offer/answer procedure in DL, or alternatively, UL between the SRC and the SRS is performed over the SR-4, or alternatively RTC-4, user plane interface.
- the SDP offer/answer procedure in DL, or alternatively, UL between the SRC and the SRS is performed over the SR-4m media- centric user plane interface over a 5GS.
- the SDP offer/answer procedure in DL, or alternatively, UL between the SRC and a dedicated signaling server interfacing with the SRS is performed over the SR-4s signaling-centric user plane interface over a 5GS.
- a 3GPP specific SDP attribute may be utilized over a 5GS implementation to convey the UUIDs determining the type and format of interaction and immersion metadata payloads over a video elementary stream as SDP attributes.
- one or more SDP attributes may be used to list one or more UUIDs for a video stream.
- a UUID may be assigned to more than one video media stream.
- the UUID format may be compliant in some implementations with the ISO-IEC-11578 ANNEX A, or SMM920220294-GR-NP alternatively, with the ISO-IEC 9834-8 format.
- the fields of the UUID listed in the ABNF description above may be thus derived in an example as Version 4 UUID of ISO- IEC 9834-8, i.e., at random.
- the H.264 media stream on port 49230 embeds in the video elementary stream UUIDs for viewport pose tracking information, user pose information and user hand tracking information
- the H.265 media stream on port 49231 embeds in the video elementary stream UUIDs for viewport pose tracking information and user pose information only.
- the SRC or alternatively, the SRS handle the user data post-decoding and expose it further to other functional blocks (e.g., XR runtime, split-rendering-aware application, SRS media processing functionality).
- This disclosure herein proposes the necessary configuration signaling to enable XR applications split rendering, with transport of interaction and immersion user data over video elementary streams, as SEI messages (H.264, H.265, H.266) or alternatively OBU metadata (AV1).
- SEI messages H.264, H.265, H.266
- OBU metadata AV1
- the solution supports multiple types of metadata based on a UUID identifying the type and format of each metadata payload to be carried over the video elementary stream.
- the proposed signaling mechanisms for the media SMM920220294-GR-NP configuration is based on three methods: application-layer signaling from the ASP to the application followed by update of the media configuration to the split-rendering client for the split-rendering media delivery session; control-plane signaling using the split- rendering application function to indicate the media session handler the media configuration followed by an update of the media configuration to the split-rendering client for the split-rendering media delivery session; user-plane signaling by SDP upon split-rendering media delivery session establishment based on a new SDP attribute; the new SDP attribute can configure for each video media flow a list of UUIDs identifying the carried metadata on the video elementary stream.
- the problem solved by this disclosure is the media configuration signaling for real-time transport of interaction and immersion multimedia data for split-rendering interactive and immersive XR applications.
- the user interaction and immersion data e.g., pose information, FoV tracking, user actions
- the XR runtime needs metadata information about the user pose and inputs that were used by the split-rendering server to render the frames. These are needed for late stage reprojection.
- real-time transport mechanisms for rendered interaction and immersion metadata and media configuration signaling are necessary for efficient display of split-rendered frames.
- the invention solves the problem by utilizing SEI messages and OBU metadata as transport mechanism for the interaction and immersion metadata of a split-rendering architecture.
- the transport relies on video coded elementary streams, thus grouping the rendered frames with their associated pose information that was used for rendering at the edge.
- the media configuration needs to include signaling that identifies the type of metadata carried over the video elementary streams, as well as its mapping to a video media flow (e.g., a 5-tuple).
- Three methods are proposed for the signaling: i). the application signaling path, ii). the control plane signaling path of a real- time communications system AF and iii). the user plane signaling based on SDP offer/answer upon RTP session establishment.
- the proposed solution is superior to the data channel approach using WebRTC SCTP stack as it benefits of the advantages of RTP/SRTP, i.e., timing, synchronization, jitter management, and reliability based on FEC. Furthermore, the proposed method inherently synchronizes the interaction and immersion metadata with the one or more video streams that are used for piggybacking. Any kind of signaling is out of scope of WebRTC. SMM920220294-GR-NP [0254] The proposed solution is better than the RTP header extension approach as it does not limit the maximum payload size for an interaction and immersion multimedia data type. Furthermore, the proposed solution is transported as part of the RTP payload and thus benefits of RTP synchronization, jitter management and reliability by means of FEC.
- RTP header extension signaling relies exclusively on SDP offer/answer procedure.
- SDP is a candidate signaling in the proposed disclosure, yet a new SDP attribute is introduced which is orthogonal to any RTP header extensions related signaling solutions.
- the proposed solution is a trade-off with respect to the approach of defining a new IETF RTP payload for interaction and immersion multimedia data. The trade-off is mainly targeted at circumventing the need of providing a full RTP payload type specification.
- the disclosure provides an embodiment that utilizes application- based signaling. The signaling of the renderer interaction and immersion metadata and subsequent mapping to the video media flows is performed by the application service provider which informs the application client of the media configuration.
- a further embodiment utilizes network-supported signaling.
- the signaling of the renderer interaction and immersion metadata and subsequent mapping to the video media flows is performed by the split rendering AF which informs the media session handler of the media configuration.
- the media session handler exposes this configuration over its APIs to the application client but also to the split-rendering client.
- a further embodiment utilizes SDP-based signaling. The signaling of the renderer interaction and immersion metadata and subsequent mapping to the video media flows is performed by means of the SDP offer/answer procedure just before the establishment of the media session.
- a method for configuration of a multimedia split-rendering content delivery session over a network comprising of: determining a media configuration comprising a mapping of one or more interaction and immersion multimedia data types to one or more video media flows, whereby the one or more interaction and immersion multimedia data types contain non-video coded data from one or more media sources; signaling the SMM920220294-GR-NP determined media configuration to a remote endpoint; establishing in part based on the signaled media configuration a multimedia split-rendering content delivery session configuration including at least one video media flow of at least one video coded elementary stream comprising the one or more mapped interaction and immersion multimedia data as non-video coded metadata payloads; utilizing the non
- the non-video coded metadata payload comprises at least two fields.
- a first field indicates a syntax and semantics identifier representation format of an interaction and immersion multimedia data type
- a second field encodes the interaction and immersion multimedia data according to the identifier representation format determined by the first field.
- the first field comprises a universally unique identifier (UUID) indication.
- UUID universally unique identifier
- Some embodiments further comprise the configuration mapping matching one or more identifier representation formats corresponding to one or more interaction and immersion multimedia data types to one or more video media flows.
- Some embodiments further comprise the signaling of the media configuration being performed based on at least one of: a signaling performed by an Application Service Provider (ASP) to a split-rendering-aware application over an application interface; a signaling between a real-time communications application function (RTC AF) and a Media Session Handler over a control plane interface (e.g., RTC-5 reference interface in a 5GS between the RTC AF and the MSH); and a signaling between a split- rendering server (SRS) and a split-rendering client (SRC) over a user plane interface (e.g., SR-4, or alternatively, RTC-4 reference interface in a 5GS between a SRS and a SRC).
- ASP Application Service Provider
- RTC AF real-time communications application function
- SRC split-rendering client
- Some embodiments further comprise the determination of the media configuration based in part on at least one of: an Application Service Provider (ASP) configuration provisioning to at least one of a real-time communications application function (RTC AF), a provisioning function, a split-rendering application function (SR AF) and a configuration function; and an application configuration of media capabilities SMM920220294-GR-NP determined semi-statically based partly on hardware capabilities of a corresponding device processing the logic of the application.
- the at least one video coded elementary stream is determined according to at least one of: H.264 video codec specification; H.265 video codec specification; H.266 video codec specification; and AV1 video codec specification.
- the non-video coded metadata payloads comprised in the at least one video coded elementary stream include at least one of: a supplemental enhancement information (SEI) message of type user data unregistered; and a metadata open bitstream unit (OBU) of type unregistered user private data.
- SEI Supplemental Enhancement Information
- OBU metadata open bitstream unit
- the interaction and immersion multimedia data types comprise one of: a user pose data representation (as for instance a timestamped 3D positional vector and quaternion representation of an XR space describing a pose object orientation up to 6DoF.
- Such a pose object may correspond to a user body component or segment, such as head, joints, hands or a combination thereof); a user gesture tracking data representation (i.e., an array of one or more hands tracked according to their pose, each hand tracking additionally consisting of an array of hand joint locations and velocities relative to a base XR space and XR runtime timestamp, for example, as per OpenXR OpenXR_EXT_hand_tracking API specification); a user body tracking data representation (e.g., a BioVision Hierarchical (BVH) encoding of the body and body segments movements and associated pose object); a user facial features tracking data representation (e.g., as an array of key points/features positions, pose, or their encoding to pre-determined facial expression classes); a set of one or more user actions (i.e., user actions and inputs to physical controllers or logic controllers defined within an XR space, for example as per OpenXR XrAction handle capturing diverse user inputs to controllers or H
- Some embodiments further comprise the split-rendering utilization of the interaction and immersion data for at least one of: a late stage reprojection at a split- rendering client (SRC) for displaying one or more rendered frames matching latest user pose and orientation information; a partial or full pre-rendering of next one or more SMM920220294-GR-NP frames at a split-rendering server (SRS); and an estimation of at a split-rendering server (SRS) of the most probable user pose and orientation associated with the expected display time of the next one or more frames to be partially or fully pre-rendered.
- SRC split- rendering client
- SRS split-rendering server
- the mapping of the interaction and immersion multimedia data types to the media flows is based on at least one of a 5-tuple indication (e.g., as a 5- tuple describing an IP flow comprising of (IP source address, IP destination address, source port, destination port, protocol identifier)), and a media flow identifier (e.g., a media description under SDP, a media flow unique name as provided in the SDP media description attribute etc.).
- a 5-tuple indication e.g., as a 5- tuple describing an IP flow comprising of (IP source address, IP destination address, source port, destination port, protocol identifier)
- a media flow identifier e.g., a media description under SDP, a media flow unique name as provided in the SDP media description attribute etc.
- the disclosure herein also provides, from a split-rendering client perspective, a method for configuration of a multimedia split-rendering content delivery session over a network, the method comprising of: receiving over the network a media configuration comprising a mapping of one or more interaction and immersion multimedia data types to one or more video media flows, whereby the one or more interaction and immersion multimedia data types contain non-video coded data from one or more media sources; controlling a video decoder to generate a set of one or more information payloads corresponding to the non-video coded payloads, whereby each of non-video coded payloads comprises of at least two fields; processing the one or more information payloads as one or more samples of interaction and immersion multimedia data generated by one or more media sources.
- the non-video coded metadata payload comprises at least two fields.
- a first field indicates a syntax and semantics identifier representation format of an interaction and immersion multimedia data type
- a second field encodes the interaction and immersion multimedia data according to the identifier representation format determined by the first field.
- the first field comprises a universally unique identifier (UUID) indication.
- UUID universally unique identifier
- Some embodiments further comprise the configuration mapping matching one or more identifier representation formats corresponding to one or more interaction and immersion multimedia data types to one or more video media flows.
- Some embodiments further comprise the receiving of the media configuration being performed based on at least one of: a signaling performed by an Application Service Provider (ASP) to a split-rendering-aware application over an application SMM920220294-GR-NP interface; a signaling between a real-time communications application function (RTC AF) and a Media Session Handler over a control plane interface (e.g., RTC-5 reference interface in a 5GS between the RTC AF and the MSH); and a signaling between a split- rendering server (SRS) and a split-rendering client (SRC) over a user plane interface (e.g., SR-4, or alternatively, RTC-4 reference interface in a 5GS between a SRS and a SRC).
- a signaling performed by an Application Service Provider (ASP) to a split-rendering-aware application over an application SMM920220294-GR-NP interface e.g., a real-time communications application function (RTC AF
- Some embodiments further comprise the media configuration being based in part on at least one of: an Application Service Provider (ASP) configuration provisioning to at least one of a real-time communications application function (RTC AF), a provisioning function, a split-rendering application function (SR AF) and a configuration function; and an application configuration of media capabilities determined semi-statically based partly on hardware capabilities of a corresponding device processing the logic of the application.
- ASP Application Service Provider
- RTC AF real-time communications application function
- SR AF split-rendering application function
- the at least one video coded elementary stream is determined according to at least one of: H.264 video codec specification; H.265 video codec specification; H.266 video codec specification; and AV1 video codec specification.
- the non-video coded metadata payloads comprised in the at least one video coded elementary stream include at least one of: a supplemental enhancement information (SEI) message of type user data unregistered; and a metadata open bitstream unit (OBU) of type unregistered user private data.
- SEI Supplemental Enhancement Information
- OBU metadata open bitstream unit
- the interaction and immersion multimedia data types comprise one of: a user pose data representation (as for instance a timestamped 3D positional vector and quaternion representation of an XR space describing a pose object orientation up to 6DoF.
- Such a pose object may correspond to a user body component or segment, such as head, joints, hands or a combination thereof); a user gesture tracking data representation (i.e., an array of one or more hands tracked according to their pose, each hand tracking additionally consisting of an array of hand joint locations and velocities relative to a base XR space and XR runtime timestamp, for example, as per OpenXR OpenXR_EXT_hand_tracking API specification); a user body tracking data representation (e.g., a BioVision Hierarchical (BVH) encoding of the body and body segments movements and associated pose object); a user facial features tracking data representation (e.g., as an array of key points/features positions, pose, or their encoding to pre-determined facial expression classes); a set of one or more user actions (i.e., user actions and inputs to physical controllers or logic controllers defined within an XR space, for example as per OpenXR XrAction handle capturing diverse user inputs to controllers SMM
- Some embodiments further comprise the split-rendering utilization of the interaction and immersion data for at least one of: a late stage reprojection at a split- rendering client (SRC) for displaying one or more rendered frames matching latest user pose and orientation information; a partial or full pre-rendering of next one or more frames at a split-rendering server (SRS); and an estimation of at a split-rendering server (SRS) of the most probable user pose and orientation associated with the expected display time of the next one or more frames to be partially or fully pre-rendered.
- SRC split- rendering client
- SRS split-rendering server
- the mapping of the interaction and immersion multimedia data types to the media flows is based on at least one of a 5-tuple indication (e.g., as a 5- tuple describing an IP flow comprising of (IP source address, IP destination address, source port, destination port, protocol identifier)), and a media flow identifier (e.g., a media description under SDP, a media flow unique name as provided in the SDP media description attribute etc.).
- a 5-tuple indication e.g., as a 5- tuple describing an IP flow comprising of (IP source address, IP destination address, source port, destination port, protocol identifier)
- a media flow identifier e.g., a media description under SDP, a media flow unique name as provided in the SDP media description attribute etc.
- the method may also be embodied in a set of instructions, stored on a computer readable medium, which when loaded into a computer processor, Digital Signal Processor (DSP) or similar, causes the processor to carry out the hereinbefore described methods.
- DSP Digital Signal Processor
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Computer Security & Cryptography (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
- Mobile Radio Communication Systems (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GR20230100211 | 2023-03-15 | ||
| PCT/EP2023/063124 WO2024088600A1 (en) | 2023-03-15 | 2023-05-16 | Split-rendering configuration for multimedia immersion and interaction data in a wireless communication system |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4595447A1 true EP4595447A1 (en) | 2025-08-06 |
Family
ID=86609486
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23727506.0A Pending EP4595447A1 (en) | 2023-03-15 | 2023-05-16 | Split-rendering configuration for multimedia immersion and interaction data in a wireless communication system |
Country Status (6)
| Country | Link |
|---|---|
| EP (1) | EP4595447A1 (en) |
| JP (1) | JP2026510623A (en) |
| KR (1) | KR20250158011A (en) |
| CN (1) | CN120303944A (en) |
| GB (1) | GB2639789A (en) |
| WO (1) | WO2024088600A1 (en) |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US12035020B2 (en) * | 2021-05-12 | 2024-07-09 | Qualcomm Incorporated | Split rendering of extended reality data over 5G networks |
-
2023
- 2023-05-16 EP EP23727506.0A patent/EP4595447A1/en active Pending
- 2023-05-16 KR KR1020257017695A patent/KR20250158011A/en active Pending
- 2023-05-16 WO PCT/EP2023/063124 patent/WO2024088600A1/en not_active Ceased
- 2023-05-16 JP JP2025526821A patent/JP2026510623A/en active Pending
- 2023-05-16 CN CN202380082281.7A patent/CN120303944A/en active Pending
- 2023-05-16 GB GB2506465.0A patent/GB2639789A/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| GB2639789A (en) | 2025-10-01 |
| GB202506465D0 (en) | 2025-06-11 |
| CN120303944A (en) | 2025-07-11 |
| JP2026510623A (en) | 2026-04-10 |
| WO2024088600A1 (en) | 2024-05-02 |
| KR20250158011A (en) | 2025-11-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12035020B2 (en) | Split rendering of extended reality data over 5G networks | |
| US11711505B2 (en) | Viewport dependent delivery methods for omnidirectional conversational video | |
| US20190104326A1 (en) | Content source description for immersive media data | |
| US20130019024A1 (en) | Wireless 3d streaming server | |
| US9674499B2 (en) | Compatible three-dimensional video communications | |
| US12489958B2 (en) | Split rendering of extended reality data over 5G networks | |
| US20260067741A1 (en) | Pdu set definition in a wireless communication network | |
| WO2024088603A1 (en) | Pdu set importance marking in qos flows in a wireless communication network | |
| WO2024088599A1 (en) | Transporting multimedia immersion and interaction data in a wireless communication system | |
| WO2024060719A1 (en) | Data transmission methods, apparatus, electronic device, and storage medium | |
| WO2024056199A1 (en) | Signaling pdu sets with application layer forward error correction in a wireless communication network | |
| WO2024088600A1 (en) | Split-rendering configuration for multimedia immersion and interaction data in a wireless communication system | |
| WO2024088609A1 (en) | Internet protocol version signaling in a wireless communication system | |
| US20240195966A1 (en) | A method, an apparatus and a computer program product for high quality regions change in omnidirectional conversational video | |
| US12610060B2 (en) | Setting PDU set importance for immersive media streams | |
| US20250254340A1 (en) | Setting PDU Set Importance for Immersive Media Streams | |
| EP4730815A1 (en) | Negotiating the level of detail in v-dmc | |
| US12375634B2 (en) | Method, an apparatus and a computer program product for spatial computing service session description for volumetric extended reality conversation | |
| EP4636543A1 (en) | Method to manage light data | |
| KR20260006549A (en) | PDU Set Marking in QoS Flows in Wireless Communication Networks | |
| AU2024217147A1 (en) | Pdu set marking in qos flows in a wireless communication network | |
| WO2024089679A1 (en) | Multimedia subprotocols over real time protocol | |
| EP4620177A1 (en) | Pdu set marking in qos flows in a wireless communication network | |
| GB2635341A (en) | Apparatus, method and computer program | |
| CN120604519A (en) | Signaling gesture information to a separate rendering server for an augmented reality communication session |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250429 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
| 17Q | First examination report despatched |
Effective date: 20260203 |