EP4736458A1 - Multi-view multiplane-imaging video streaming - Google Patents

Multi-view multiplane-imaging video streaming

Info

Publication number
EP4736458A1
EP4736458A1 EP24743159.6A EP24743159A EP4736458A1 EP 4736458 A1 EP4736458 A1 EP 4736458A1 EP 24743159 A EP24743159 A EP 24743159A EP 4736458 A1 EP4736458 A1 EP 4736458A1
Authority
EP
European Patent Office
Prior art keywords
video
mpi
sequence
atlas
bitstream
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24743159.6A
Other languages
German (de)
French (fr)
Inventor
Sejin OH
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Dolby Laboratories Licensing Corp
Original Assignee
Dolby Laboratories Licensing Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Dolby Laboratories Licensing Corp filed Critical Dolby Laboratories Licensing Corp
Publication of EP4736458A1 publication Critical patent/EP4736458A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/21Server components or server architectures
    • H04N21/218Source of audio or video content, e.g. local disk arrays
    • H04N21/21805Source of audio or video content, e.g. local disk arrays enabling multiple viewpoints, e.g. using a plurality of cameras
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/597Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding specially adapted for multi-view video sequence encoding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/70Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by syntax aspects related to video coding, e.g. related to compression standards
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/234Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
    • H04N21/2343Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements
    • H04N21/23439Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements for generating different versions
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/235Processing of additional data, e.g. scrambling of additional data or processing content descriptors
    • H04N21/2353Processing of additional data, e.g. scrambling of additional data or processing content descriptors specifically adapted to content descriptors, e.g. coding, compressing or processing of metadata
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/60Network structure or processes for video distribution between server and client or between remote clients; Control signalling between clients, server and network components; Transmission of management data between server and client, e.g. sending from server to client commands for recording incoming content stream; Communication details between server and client 
    • H04N21/65Transmission of management data between client and server
    • H04N21/658Transmission by the client directed to the server
    • H04N21/6587Control parameters, e.g. trick play commands, viewpoint selection
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/81Monomedia components thereof
    • H04N21/816Monomedia components thereof involving special video data, e.g 3D video
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/84Generation or processing of descriptive data, e.g. content descriptors
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/845Structuring of content, e.g. decomposing content into time segments
    • H04N21/8456Structuring of content, e.g. decomposing content into time segments by decomposing the content in the time domain, e.g. in time segments

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Library & Information Science (AREA)
  • Databases & Information Systems (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)
  • Testing, Inspecting, Measuring Of Stereoscopic Televisions And Televisions (AREA)

Abstract

Methods and apparatus for multiplane-imaging (MPI) video streaming. According to an example embodiment, a method for streaming an MPI video includes generating a sequence of video frames, each of the video frames including a respective plurality of patches representing texture and transparency layers of one or more multiplane images of the MPI video and applying video compression to the sequence of video frames to generate a video sub-stream. The method also includes generating a sequence of representations of atlas frames corresponding to the sequence of video frames to specify at least a packing arrangement of the patches and applying compression to the sequence of representations to generate an atlas sub-stream. The method further includes multiplexing the video sub-stream and the atlas sub-stream to generate a first coded bitstream encoding at least a portion of the MPI video.

Description

Docket No. D23136WO01 MULTI-VIEW MULTIPLANE-IMAGING VIDEO STREAMING 1. Cross-Reference to Related Applications [0001] This patent application claims the benefit of priority to the following applications: U.S. Provisional Patent Application No.63/556,461 filed February 22, 2024, U.S. Provisional Patent Application No.63/588,337 filed October 6, 2023, and U.S. Provisional Patent Application No.63/510,571 filed June 27, 2023, each of which is hereby incorporated by reference in their entireties. 2. Field of the Disclosure [0002] Various example embodiments relate generally to multiplane imaging (MPI) and, more specifically but not exclusively, to transmission of multiplane images. 3. Background [0003] Multiplane images embody a relatively new approach to storing volumetric content. MPI can be used to render both still images and video and represents a three- dimensional (3D) scene within a view frustum using, e.g., 8, 16, or 32 planes of texture and transparency (alpha) information per camera. Example applications of MPI include computer vision and graphics, image editing, photo animation, robotics, and virtual reality. BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS [0004] Example embodiments disclosed herein provide formats for coding, storage, and delivery of multi-view MPI video and a corresponding MPI video system. Some examples use the ISO Base Media File Format (BMFF) for storage and MPEG Dynamic Adaptive Streaming over HTTP (DASH) for streaming of MPI video to provide immersive volumetric video experience. Some examples enable seamless interactive view switching for multi-view MPI video streaming using modifications to available technologies and infrastructure, e.g., by defining new tools designed to fill the technological gap for the above-stated purposes. At least some of the disclosed solutions can be deployed in a relatively short time, after modifications of pertinent existing solutions are implemented in accordance with various embodiments disclosed herein. [0005] According to an example embodiment, provided is a method for streaming an MPI video, the method comprising: generating a sequence of video frames, each of the video frames including a respective plurality of patches representing texture and transparency layers of one or more multiplane images of the MPI video; applying video compression to the sequence of video frames to generate a video sub-stream; generating a sequence of representations of atlas Docket No. D23136WO01 frames corresponding to the sequence of video frames to specify at least a packing arrangement of the patches; applying compression to the sequence of representations to generate an atlas sub- stream; and multiplexing the video sub-stream and the atlas sub-stream to generate a first coded bitstream encoding at least a portion of the MPI video. [0006] According to another example embodiment, provided is a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the above method. [0007] According to yet another example embodiment, provided is An apparatus for streaming an MPI video, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: generate a sequence of video frames, each of the video frames including a respective plurality of patches representing texture and transparency layers of one or more multiplane images of the MPI video; apply video compression to the sequence of video frames to generate a video sub-stream; generate a sequence of representations of atlas frames corresponding to the sequence of video frames to specify at least a packing arrangement of the patches; apply compression to the sequence of representations to generate an atlas sub-stream; and multiplex the video sub-stream and the atlas sub-stream to generate a first coded bitstream encoding at least a portion of the MPI video. BRIEF DESCRIPTION OF THE DRAWINGS [0008] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which: [0009] FIG.1 depicts an example process for a video/image delivery pipeline. [0010] FIG.2 pictorially illustrates a 3D-scene representation using a multiplane image according to an embodiment. [0011] FIG.3 pictorially illustrates a process of generating a novel view of a 3D scene according to one example. [0012] FIG.4 is a block diagram illustrating a change of the set of active views over time according to one example. [0013] FIG.5 is a block diagram illustrating an MPI video system that can be used in the delivery pipeline of FIG.1 according to an embodiment. [0014] FIG.6 is a block diagram illustrating an MPI encoder that can be used in the MPI video system of FIG.5 according to an embodiment. Docket No. D23136WO01 [0015] FIG.7 is a block diagram illustrating an example structure of an output V3C bitstream that can be generated by the MPI encoder of FIG.6 according to an embodiment. [0016] FIGS.8A-8F are diagrams illustrating examples of spatial packing of texture and transparency maps of MPI layers into video frames that can be implemented in the MPI video system of FIG.5 according to some embodiments. [0017] FIG.9 is a block diagram illustrating single-track encapsulation of a V3C bitstream that can be implemented in the MPI video system of FIG.5 according to some examples. [0018] FIGS.10A-10B are block diagrams illustrating multiple-track encapsulation of a V3C bitstream that can be implemented in the MPI video system of FIG.5 according to some examples. [0019] FIG.11 is a block diagram illustrating further details of the multiple-track encapsulation corresponding to the example illustrated in FIG.10A. [0020] FIG.12 is a block diagram illustrating further details of the multiple-track encapsulation corresponding to the example illustrated in FIG.10B. [0021] FIG.13 is a block diagram illustrating single-track encapsulation of a 2D video stream (MPI bitstream) that can be implemented in the MPI video system of FIG.5 according to some examples. [0022] FIG.14 is a block diagram illustrates a DASH configuration for grouping atlas and packed video components that can be implemented in the MPI video system of FIG.5 according to some examples. [0023] FIG.15 is a flowchart illustrating client-server communications and the corresponding processing operations performed at the client that can be implemented in the MPI video system of FIG.5 according to various examples. [0024] FIG.16 is a block diagram illustrating a DASH configuration that can be used in the MPI video system of FIG.5 for the support of seamless switching among MPI views according to some examples. [0025] FIG.17 is a block diagram of an example computing device, one or more instances of which can be used to implement the MPI video system of FIG.5 according to various examples. DETAILED DESCRIPTION Example Video/Image Delivery Pipeline Docket No. D23136WO01 [0026] FIG.1 depicts an example process of a video delivery pipeline (100), showing various stages from video/image capture to video/image-content display according to an embodiment. A sequence of video/image frames (102) may be captured or generated using an image-generation block (105). The frames (102) may be digitally captured (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video and/or image data (107). Alternatively, the frames (102) may be captured on film by a film camera. Then, the film may be translated into a digital format to provide the video/image data (107). In some examples, the image-generation block (105) includes generating an MPI image or video. [0027] In a production phase (110), the data (107) may be edited to provide a video/image production stream (112). The data of the video/image production stream (112) may be provided to a processor (or one or more processors, such as a central processing unit, CPU) at a post-production block (115) for post-production editing. The post-production editing of the block (115) may include, e.g., adjusting or modifying colors or brightness in particular areas of an image to enhance the image quality or achieve a particular appearance for the image in accordance with the video creator’s creative intent. This part of post-production editing is sometimes referred to as “color timing” or “color grading.” Other editing (e.g., scene selection and sequencing, image cropping, addition of computer-generated visual special effects, removal of artifacts, etc.) may be performed at the block (115) to yield a “final” version (117) of the production for distribution. In some examples, operations performed at the block (115) include enhancing texture and/or alpha channels in multiplane images/video. During the post-production editing (115), video and/or images may be viewed on a reference display (125). [0028] Following the post-production (115), the data of the final version (117) may be delivered to a coding block (120) for being further delivered downstream to decoding and playback devices, such as television sets, set-top boxes, movie theaters, and the like. In some embodiments, the coding block (120) may include audio and video encoders, such as those defined by the ATSC, DVB, DVD, Blu-Ray, and other delivery formats, to generate a coded bitstream (122). In a receiver, the coded bitstream (122) is decoded by a decoding unit (130) to generate a corresponding decoded signal (132) representing a copy or a close approximation of the signal (117). The receiver may be attached to a target display (140) that may have somewhat or completely different characteristics than the reference display (125). In such cases, a display management (DM) block (135) may be used to map the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137). Depending on the embodiment, the decoding unit (130) and display management block (135) may include individual processors or may be based on a single integrated processing unit. Docket No. D23136WO01 [0029] A codec used in the coding block (120) and/or the decoding block (130) enables video/image data processing and compression/decompression. The compression is used in the coding block (120) to make the corresponding file(s) or stream(s) smaller. The decoding process carried out by the decoding block (130) typically includes decompressing the received video/image data file(s) or streams(s) into a form usable for playback and/or further editing. Example coding/decoding operations that can be used in the coding block (120) and the decoding unit (130) according to various embodiments are described in more details below. Multiplane Imaging [0030] A multiplane image comprises multiple image planes, with each of the image planes being a “snapshot” of the 3D scene at a certain depth with respect to the camera position. Information stored in each plane includes the texture information (e.g., represented by the R, G, B values) and transparency information (e.g., represented by the alpha (A) values). Herein, the acronyms R, G, B stand for red, green, and blue, respectively. In some examples, the three texture components can o be (Y, Cb, Cr), or (I, Ct, Cp), or another functionally similar set of values. There are different ways in which a multiplane image can be generated. For example, two or more input images from two or more cameras located at different known viewpoints can be co-processed to generate a corresponding multiplane image. Alternatively, single-view synthesis of a multiplane image can be performed using a source image captured by a single camera. [0031] FIG.2 pictorially illustrates a 3D scene representation using a multiplane image (200) according to an embodiment. The multiplane image (200) has D planes or layers (P0, P1, …, P(D−1)), where D is an integer greater than one. Typically, the planes (layers) are indexed such that the most remote layer, from the reference camera position (RCP), is indexed as the 0-th layer and is at a distance (or depth) d0 from the RCP along the Z dimension of the 3D scene. The index is incremented by one for each next layer located closer to the RCP. The plane (layer) that is the closest to the RCP has the index value (D−1) and is at a distance (or depth) dD1 from the RCP along the Z dimension. Each of the planes (P0, P1, …, P(D−1)) is orthogonal to a base plane (202) which is parallel to the XZ-coordinate plane. The RCP is at a vertical height h above the base plane (202). The XYZ triad shown in FIG.2 indicates the general orientation of the multiplane image (200) and the planes (P0, P1, …, P(D−1)) with respect to the X, Y, and Z dimensions of the 3D scene. In various examples, the number D can be 32, 16, 8, or any other suitable integer greater than one. [0032] Let us denote the color component (e.g., RGB) value for the ith layer at camera location s as ^^^^( ^^^^) ^^^^ , with the lateral size of the layer being H×W, where H is the height (Y Docket No. D23136WO01 dimension) and W is the width (X dimension) of the layer. The pixel value at location (x, y) for the color channel c is represented as ^^^^( ^^^^) th ( ^^^^) ^^^^ ( ^^^^, ^^^^, ^^^^). The α value for the i layer is ^^^^ ^^^^ . The pixel value (x, y) in the alpha layer is represented as ^^^^( ^^^^) th ^^^^ ( ^^^^, ^^^^). The depth distance between the i layer to the reference camera position is ^^^^ ^^^^. The image from the original reference view (without the camera moving) is denoted as ^^^^, with the texture pixel value being ^^^^). A still MPI image for the camera location s can therefore be represented as: It is straightforward to extend this still MPI image representation to a video representation, provided that that the camera position s is kept static overtime. This video representation is given by Eq. (2): where t denotes time. [0033] As already indicated above, a multiplane image, such as the multiplane image (200), can be generated using single-view synthesis from a single source image ^^^^ or using multiple-view synthesis from two or more source images. Such syntheses may be performed, e.g., during the production phase (110). The corresponding MPI synthesis algorithm(s) may typically output the multiplane image (200) containing XYZ-resolved pixel values in the form {( ^^^^ ^^^^, ^^^^ ^^^^) for i=0, …, D−1}. [0034] By processing the multiplane image (200) represented by {( ^^^^ ^^^^, ^^^^ ^^^^) for i=0, …, D−1}, an MPI-rendering algorithm can generate a viewable image corresponding to the RCP or to a new virtual camera position that is different from the RCP. An example MPI-rendering algorithm (often referred to as the “MPI viewer”) that can be used for this purpose may include the steps of warping and compositing. Other suitable MPI viewers may also be used. The rendered multiplane image (200) can be viewed, e.g., on the reference display (125). [0035] During the warping step of the MPI-rendering algorithm, each layer ( ^^^^ ^^^^, ^^^^ ^^^^) of the multiplane image (200) may be warped from the RCP viewpoint position ( ^^^^ ^^^^) to a new viewpoint position ( ^^^^ ^^^^), e.g., as follows: ^^^^ ^ ^^ ^ ^^ ^^ = ^^^^ ^^^^ ^^^^, ^^^^ ^^^^( ^^^^ ^^^^ ^^^^ , ^^^^ ^^^^) (3) where ^^^^ ^^^^ ^^^^, ^^^^ ^^^^() is the warping function; and σ is the consistent scale (to minimize error). In an example embodiment, the warping function ^^^^ ^^^^ ^^^^, ^^^^ ^^^^() can be expressed as follows: (5) Docket No. D23136WO01 where ^^^^ ^^^^ = ( ^^^^ ^^^^, ^^^^ ^^^^) and ^^^^ ^^^^ = ( ^^^^ ^^^^ , ^^^^ ^^^^). Through (5), each pixel location ( ^^^^ ^^^^ , ^^^^ ^^^^) on the target view of a certain MPI plane can be mapped to its respective pixel location ( ^^^^ ^^^^, ^^^^ ^^^^) on the source view. The functions ^^^^ ^^^^ and ^^^^ ^^^^ represent the intrinsic camera model for the reference view and the target view, respectively. The functions R and t represent the extrinsic camera model for rotation and translation, respectively. n denotes the normal vector [001]T. a denotes the distance to a plane that is fronto-parallel to the source camera at depth ^^^^ ^^^^ ^^^^. [0036] During the compositing step of the MPI-rendering algorithm, a new viewable image ^^^^ ^^^^ can be generated, e.g., using processing operations corresponding to the following equations: = ∑ ^ ^^ ^ ^^ ^ =^− 01 ^^^^ ^ ^^ ^ ^^ ^^ ^^^^ ^ ^^ ^ ^^ ^^ (6) where the weights ^^^^ ^ ^^ ^ ^^ ^^ are expressed as: ^^^^ ^ ^^ ^ ^^ ^^ = ( ^^^^ ^ ^^ ^ ^^ ^^ ∏ ^ ^^ ^ ^^ ^^ = ^^^ 1 ^+1 (1 − ^^^^ ^ ^ ^^ ^^ ^^ ) ) (7) The disparity map ^^^^ ^^^^ corresponding to the source view can be computed as: ^^^^ ^^^^ = ∑ ^ ^^ ^ ^^ ^ =^− 01 ^^^^ −1 ^^^^ ^^^^ ^ ^ ^^ ^^ ^^ (8) where the weights ^^^^ ^ ^ ^^ ^^ ^^ are expressed as: The MPI-rendering algorithm can also be used to generate the viewable image ^^^^ ^^^^ corresponding to the RCP. In this case, the warping step is omitted, and the image ^^^^ ^^^^ is computed as: ^^^^ ^^^^ = ∑ ^ ^^ ^ ^^ ^ =^− 01 ^^^^ ^ ^ ^^ ^^ ^^ ^^^^ ^ ^ ^^ ^^ ^^ (10) [0037] In the single camera transmission scenario, only one MPI is fed through a bitstream. A goal for this situation is to optimally merge the layers of the original MPI such that the quality of this MPI after local warping is preserved. In the multiple camera transmission scenario, multiple MPIs captured in different camera positions are encoded in the compressed bitstream. The information in these MPIs is jointly used to generate global novel views for positions located between the original camera positions. There also can be a scenario where information from multiple cameras can be used jointly to generate a single MPI to be transmitted. For transmissions of MPI video, the multiple camera transmission scenario is typically used, e.g., as explained below. [0038] FIG.3 pictorially illustrates a process of generating a novel view of a 3D scene (302) according to one example. In the example shown, the 3D scene (302) is captured using forty-two RCPs (1, 2, …, 42). The novel view that is being generated corresponds to a camera position (50). The four closest RCPs to the camera position (50) are the RCPs (11, 12, 18, 19). The corresponding multiplane images are multiplane images (20011, 20012, 20018, 20019). A multiplane image (20050) corresponding to the camera position (50) is generated by Docket No. D23136WO01 correspondingly warping the multiplane images (20011, 20012, 20018, 20019) and then merging the resulting warped multiplane images. Finally, a viewable image (312) of the 3D scene (302) is generated by applying the compositing step (310) of the MPI-rendering algorithm to the multiplane image (20050). [0039] In general, a 3D scene, such as the 3D scene (302) may be captured using any suitably selected number of RCPs. The locations of such RCPs can also be variously selected, e.g., based on the creative intent. In typical practical example, when a novel view, such as the viewable image (312) is rendered, only several neighboring RCPs are used for the rendering. Hereafter, such neighboring views are referred to as the “active views.” In the example illustrated in FIG.3, the number of active views is four. In other examples, a different (from four) number of active views may similarly be used. As such, the number of active views is a selectable parameter. For illustration purposes and without any implied limitations, example embodiments are described herein below in reference to four active views. In some examples, the set of active views may change over time when the camera position (50) moves. In some examples, the number of active views may change over time when the camera position (50) moves. [0040] FIG.4 is a block diagram illustrating a change of the set of active views over time according to one example. In the example shown, a 3D scene is captured using a rectangular array of forty RCPs arranged in five rows and eight columns. A dashed arrow (402) represents a movement trajectory of the novel camera position (50) during the time interval starting at the time t0 and ending at the time t1 for virtual view synthesis. At the time t0, the set of active views includes the four views encompassed by the dashed box (410). At the time t1, the set of active views includes the four views encompassed by the dotted box (420). At the time tn (where t0< tn < t1), the set of active views changes from being the set in the dashed box (410) to being the set in the dotted box (420). [0041] Various embodiments disclosed herein are directed at providing an end-to-end interactive multi-view MPI video streaming system including: • Components located at the server side and content storage formats used therein; • Components supporting through-the-network distribution protocols; and • Components supporting the end client-side interactive playback solutions. Some embodiments use the MPEG-DASH and ISO-BMFF standards and/or their extensions and additionally provide new tools that support new functionalities. One goal of such embodiments is to enable seamless interactive view switching for multi-view MPI video streaming by building up on and expanding certain legacy technologies/infrastructures to enable expedited and/or cost- Docket No. D23136WO01 efficient deployment of the interactive immersive volumetric video experience. Some embodiments may benefit from at least some features disclosed in one or more of the following documents: (1) ISO/IEC 14496-12:2020, Information technology -- Coding of audio-visual objects -- Part 12: ISO base media file format; (2) ISO/IEC 14496-15:2021, Information technology -- Coding of audio-visual objects -- Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format; (3) ISO/IEC 23000-19:2018, Information technology — Multimedia application format (MPEG-A) -- Part 19: Common media application format (CMAF) for segmented media, and (4) ISO/IEC 23009-1(2019), Information Technology -- Dynamic Adaptive Streaming over HTTP (DASH) -- Part 1: Media Presentation Description and Segment Formats, . Each of these documents is incorporated herein by reference in its entirety. [0042] For the storage format, some embodiments build on the ISO base media file format, commonly referred to as ISO BMFF, (ISO/IEC 14496-12), and its extension ISO/IEC 14496-15, which provides carriage of the network abstraction layer (NAL) unit structured video, such as the AVC, HEVC, VVC, and the like. Herein, the acronym ISO stands for the International Organization for Standardization. For the data transport, some embodiments build on Dynamic adaptive streaming over HTTP (DASH) (ISO/IEC 230090-1). In some examples, a similar approach can be applied to the Common Media Application Format (CMAF). Note that the disclosed video streaming system can also support other multi-view coding solutions, such as the MVC, SHVC, and multi-layer VVC. For illustration purposes and without any implied limitations, example embodiments are described in reference to the HEVC codec. Based on the provided description, a person of ordinary skill in the pertinent art will be able to make and use other embodiments compatible with other suitable codecs without any undue experimentation. Transmission of Coded MPI Videos [0043] FIG.5 is a block diagram illustrating an MPI video system (500) that can be used in the delivery pipeline (100) according to various embodiments. In different embodiments, the system (500) can be configured for live and/or on-demand use cases. The system (500) includes an MPI video source (502) in which a scene is captured by a set of cameras, and the scene acquisition results in an MPI video (504). Each MPI image of the MPI video (504) includes fronto-parallel planes of the corresponding scene sampled at different depths from the reference coordinate frame. Information stored in each plane include the texture and opacity/transparency, e.g., as described above. [0044] An MPI encoder (506) operates to encode the MPI video (504) into a coded bitstream (508), e.g., as described in more detail below. In some examples, the MPI encoder Docket No. D23136WO01 (506) uses a packing process that causes the texture and alpha maps of an MPI representation to be spatially packed and/or temporally interleaved. The MPI encoder (506) further operates to code the resulting packed video with a two-dimensional (2D) codec, such as AVC, HEVC, VVC, or the like. In some examples, the MPI representation can be coded with the MPEG immersive video coding standard. An encapsulation module (514) operates to encapsulate the coded bitstream (508) into a media file for file playback (516) or into a sequence of segments for streaming (518) according to a selected media container file format. In some examples, the media container file format is the above-mentioned ISO BMFF. The encapsulation module (514) also includes metadata (512) into its output (516 or 518). The metadata (512) are generated with a processing/authoring module (510) and may specify how the texture and alpha maps are packed into video frames carried into the media file or segment. The segments are delivered to a player (590) using a delivery mechanism (520). The player (590) includes a decapsulation module (524) configured to perform inverse operations to the operations of the encapsulation module (514). [0045] To enable the player (590) to request specific video segments among available video segments, corresponding metadata are signaled in a streaming manifest file that contains sufficient information needed for the player (590) to download and present a given piece of content. In some examples, the streaming manifest file is the Media Presentation Description (MPD) of the MPEG-DASH standard. The player (590) receives the streaming manifest file and determines which video segments are appropriate for the current configuration/situation, such as the viewing position and orientation, device capability, bandwidth, etc. Then the player (590) operates to retrieve the video segments accordingly. [0046] The decapsulation module (524) operates to process the video file or segments and reconstructs the coded bitstream (508). The decapsulation module (524) further operates to parse the metadata (512) and extract the packing information (of how the texture and alpha maps are packed into video frames at the encoder side). In some examples, the bitstream of one MPI representation may be carried in multiple tracks. An MPI decoder (528) operates to decode the coded bitstream(s) (508) into multiple frames, and a reconstructed signal (530) carrying the MPI representation is generated in a reconstruction module (532) from the corresponding decoded video signal(s) (530) based on the MPI video related information and texture and alpha maps packing metadata. An output viewport (538) of the MPI representation is rendered by a rendering module (536) and displayed on a display device (540) according to the viewport (538) including the viewing position and orientation. Docket No. D23136WO01 [0047] For the multi-view MPI experience, the source (502) generates multi-view MPI representations including a set of MPI representations, with each MPI representation being from a different respective reference coordinate frame. Every single-view MPI representation is independently coded as described below, encapsulated into a media file for the file playback (516) or into a sequence of an initialization segment and media segments for the streaming (518). As the user changes the viewport (538), including the viewing position and/or orientation, the player (590) needs to dynamically switch to appropriate MPI bitstreams accordingly. That is, the player (590) operates to determine the video segments to be received for the relevant MPI representation according to the viewport, dynamically request the segments and decodes the received segments accordingly. Then, the output MPI is reconstructed, and the output viewport of the MPI representation is displayed which fits to the changed viewport. In the example shown, these functions of the player (590) are implemented with a viewport tracking module (550). [0048] As already described above, an MPI representation includes a set of planar layers corresponding to different respective depths from a given reference point of view. Each layer has a texture (color) and transparency channels obtained by projecting the part of the 3D scene contained around the layer location on the same reference camera, which is positioned at the given reference point of view. In different examples, the sequence of texture and transparency maps of the layers from the MPI representation can be encoded using either MPEG Immersive Video (MIV) or a 2D video codec, e.g., as described in more detail below. [0049] The MIV standard is directed, inter alia, to compression of immersive video content, also known as volumetric video. The MIV standard enables storage and distribution of immersive video content over existing and future networks, for playback with 6 degrees of freedom (6 DoF) of view position and orientation within a limited viewing space and with different fields of view depending on the capture setup. In some examples, the camera rig may be a flat setup of co-aligned cameras, disposed linearly to allow significant lateral displacement or oriented divergently for a large field of view adapted to a viewing with a head mounted display (HMD). [0050] The MIV standard is part of ISO/IEC 23090 MPEG-I, a collection of standards to digitally represent immersive media. The MIV standard is also aligned with the Video-based Point Cloud Compression (V-PCC) MPEG standard (ISO/IEC 23090-5). As a result, the MIV standard references the common part of ISO/IEC 23090-5 (2nd edition) Visual Volumetric Video-based Coding (V3C). V3C provides extension mechanisms for V‑PCC and MIV. SC Docket No. D23136WO01 29/WG 03 MPEG Systems also developed a systems standard based on ISO BMFF for the carriage of V3C data (ISO/IEC 23090-10). [0051] FIG.6 is a block diagram illustrating the MPI encoder (506) according to an embodiment. In this embodiment, the MPI encoder (506) is configured to encode the MPI video (504) using the MIV standard. The resulting output is a V3C bitstream (672), which is an example of the MPI bitstream (508) (also see FIG.5). [0052] In the example shown, the MPI encoder (506) includes a patch generation and packing module (610), a texture video generation module (630), and a transparency video generation module (640). The modules (630, 640) operate to process the MPI representation input (504) to generate sequences (632, 642) of texture atlases and the transparency atlases. The patch generation and packing module (610) generates associated atlas data (612) that are subsequently used by the player (590) to reconstruct the texture and transparency maps of the MPI layers. A frame packing module (650) operates to combine the texture atlases (632) and the transparency atlases (642) into packed video data to generate a sequence (652) of video frames including the packed video data. A video compression module (660) operates to encode the video sequence (652) using a suitable video coding method, thereby generating a corresponding compressed video sub-stream (662). A patch sequence compression module (620) operates to encode the atlas data (612) into a compressed atlas-data sub-stream (620). A multiplexer (670) operates to multiplex the compressed atlas-data sub-stream (620) and the compressed video sub- stream (662), thereby generating the output V3C bitstream (672). [0053] FIG.7 is a block diagram illustrating an example structure of the V3C bitstream (672) according to an embodiment. The MIV specifications define a set of extensions and profiles in reference to V3C. For examplee, MIV supports multiple atlases and view parameters. For each atlas, geometry, occupancy, and/or attribute, information is carried in video sub- bitstreams in parallel with an atlas sub-bitstream that includes the patch parameters. View parameters (e.g., camera extrinsics, intrinsics, resolution, etc.) are carried in a common atlas sub- bitstream that is shared between the atlases. [0054] The MIV standard also defines a set of profiles to be adapted to different grades of bandwidth or decoding resources at the client side, e.g., at the player (590). Beside the Main profile based on geometry with embedded occupancy and texture and often referred to as MVD (standing for multiple video + depth), another Extended profile possibly separates occupancy from geometry and may activate transparency attribute with a Restricted Geometry sub-profile allowing for the MPI format delivery. Additionally, a Geometry Absent profile allows for generation of geometry information at the client side. Docket No. D23136WO01 [0055] In the example shown in FIG.7, the V3C bitstream (672) includes a plurality of V3C units (702, 704). The first V3C unit (702) includes a V3C parameter set (VPS) (712). The VPS (712) is configured to announce the presence of sub-bitstreams, allowing the corresponding decoder to initialize sub-decoders for all present atlas and video sub-bitstreams. Subsequent V3C units (704) carry an atlas sub-stream (714) and a video sub-stream (716). The atlas sub- stream (714) contains common atlas data (CAD), atlas data (AD). The video sub-stream (716) contains packed video data (PVD). The AD unit contains an atlas sub-bitstream, which is a network abstraction layer (NAL) unit stream, but instead of video frames there is a NAL unit referred to as the atlas tile layer (ATL) that carries a list of patch data. The CAD unit also contains common information that applies to all atlases, such as a view parameter list, including extrinsic and intrinsic view parameters. The PVD unit contains the packed video data of texture and transparency per atlas. [0056] FIGS.8A-8F are diagrams illustrating examples of spatial packing of texture and transparency maps of MPI layers into video frames that can be implemented with the MPI encoder (506) according to various embodiments. In these examples, the MPI encoder (506) is configured to use a 2D video codec. Each of FIGS.8A-8F shows a respective single video frame. In different examples, the texture and transparency maps of layers from a same MPI representation can be spatially packed into the same frame or can be temporally interleaved in various arrangements. For example, FIGS.8A and 8D show arrangements in which all texture maps of layers are grouped together in a first contiguous block, and all are transparency maps of layers are grouped together in a second contiguous block. The two contiguous blocks can then be placed side by side as in FIG.8A or top to bottom as in FIG.8D. FIGS.8B, 8C, and 8E illustrate examples in which texture and transparency maps of layers are arranged in alternating rows or columns within the video frame. FIG.4F illustrates an example in which maps of different layers are packed into regions of the video frame having different sizes, i.e., the maps can have different resolutions. In different examples, the employed 2D video codec can be an HEVC codec or a VVC codec. One or more SEI messages may contain MPI video information, e.g., including the number of MPI layers, the depth of layers, and an identification of the specific texture and transparency map arrangement, such as one of the arrangements illustrated in FIGS. 8A-8F. Herein, the acronym SEI stands for supplemental enhancement information. [0057] To assist the player (590) selecting which tracks are to be fetched from the video file or which segment(s) to download through the channel, suitable metadata may be provided. In some examples, the metadata specify: Docket No. D23136WO01 • The structure of the entire MPI representation, such as whether it is a single-view or multi- view MPI representation. • In cases of a multi-view MPI (for multiple camera setup), o Information corresponding to each reference camera view, such as: ^ Reference camera view identifier; ^ Extrinsic camera and intrinsic camera information; o Neighboring reference views to enable the player (590) to identify the views to be prefetched. • MPI video information, such as: o The number of MPI layers; o The depth information of MPI layers including whether the MPI has constant depth increment among the layers; o Packing information of the texture and transparency maps of the MPI layers; o MPI video decoding specification, including identification of the codec and profile, and level information; o Information specific to MPI post-decoder operations. [0058] At the player (590), the below-listed inputs can be used to determine which track(s) is (are) fetched from the video file or which representation(s) are to be requested. Using this information, the player (590) operates to fetch or request the desired tracks or representations so that the end user can continue consuming the MPI seamlessly. • New novel viewport, including viewing position and orientation, and the field of view: o From the user input device (eye tracking, touch screen, sensors, etc.), we know the users’ new selected viewport including viewing position and orientation, vertical and horizontal field of view. • Device characteristics: o Video decoding capability including the support codec. o Supported post-processing/rendering features, such as a more advanced algorithm to allow smooth playback when the end-user makes a large change in the novel viewing pose, e.g., configured to request more surrounding views (perhaps, at a lower bit rate for bandwidth use efficiency). Docket No. D23136WO01 • Network conditions: According to the experienced network conditions, such as bandwidth, packet loss, packet delay, the player (590) can determine which bit rate version (representation) is to be downloaded. • Current time: in some examples, we know how many frames are left in the current segment, and which segment we need to request from the server/source side. MPI Video Storage and Metadata Signaling Using ISO BMFF [0059] When an MPI representation is encoded as a V3C bitstream as described above, the V3C bitstream can be stored into a single track or multiple tracks within a file. When an MPI representation is encoded as a 2D video bitstream, it can be stored into tracks defined in ISO/IEC 14496-15. At least some of the metadata described above are signaled in tracks to enable the player (590) to select and to fetch the pertinent bitstream(s) from a file appropriately. Below, we first describe, in reference to FIGS.9-12, example embodiments for storage and metadata signaling of a V3C bitstream (518) carrying an MPI video. We then describe, in reference to FIG.13, example embodiments for storage and metadata signaling of a 2D video bitstream (518) carrying an MPI video. [0060] This portion describes encapsulation of a single V3C bitstream (518), which includes the coded MPI bitstream, in tracks. Only one of the below encapsulations may be used at the same time. • Single track encapsulation, where one track carries the entire V3C bitstream; • Multiple track encapsulation, where one or two of the multiple tracks carry the atlas sub- stream (714) and the other ones of the multiple tracks carry the packed video sub-stream (716). [0061] The box types described herein are listed in bold in Table 1. Table 1: Box types described herein (indicated in bold) and their relation to conventional boxes (not explicitly described in this specification) Docket No. D23136WO01 [0062] FIG.9 is a block diagram illustrating single-track encapsulation of the V3C bitstream (672) according to some examples. For single-track encapsulation, the V3C bitstream (672) is directly stored as a V3C bitstream track (902). The V3C unit headers are kept in the bitstream (902) as is. Each sample (920) in the V3C bitstream track (902) include respective sub-samples that belong to the sample presentation time, and each sub-sample contains a respective V3C unit containing one of common atlas data (CAD), atlas data (AD), and packed video data (PVD). [0063] A movie header (910) of the V3C bitstream track (902) contains the V3C bitstream’s decoding specific information and sub-sample information. As such, the V3C bitstream track (902) contains 2D video decoder information for configuration and initialization of the 2D video decoder. The header (910) also contains oneV3CConfigurationBox (912), which includes V3C bitstream’s decoding specific information (e.g., parameter sets and SEI messages). To enable support for the sub-sample level access, the V3C bitstream track (902) also contains one SubSampleInformationBox (916), which lists the track sub- samples. The 32-bit unit header of the V3C unit, which represents the sub-sample, is copied to the codec_specific_parameters field of the sub-sample entry in the SubSampleInformationBox (916). The V3C unit type of each sub-sample is identified by parsing the codec_specific_parameters field of the sub-sample entry in the SubSampleInformationBox (916). [0064] In some examples, the V3CConfigurationBox (912) is defined as follows: class V3CConfigurationBox extends FullBox('v3cC', version = 0, 0) { Docket No. D23136WO01 V3CDecoderConfigurationRecord(); } aligned(8) class V3CDecoderConfigurationRecord { // version 0 unsigned int(3) unit_size_precision_bytes_minus1; unsigned int(5) num_of_v3c_parameter_sets; for (int i=0; i < num_of_v3c_parameter_sets; i++) { unsigned int(16) v3c_parameter_set_length; // v3c_unit() as defined in ISO/IEC FDIS 23090-5 v3c_unit v3c_parameter_set(v3c_parameter_set_length); } unsigned int(8) num_of_setup_unit_arrays; for (int j=0; j < num_of_setup_unit_arrays; j++) { unsigned int(1) array_completeness; bit(1) reserved = 0; unsigned int(6) nal_unit_type; unsigned int(8) num_nal_units; for (int i=0; i < num_nal_units; i++) { unsigned int(16) setup_unit_length; // nal_unit(size) as defined in ISO/IEC FDIS 23090-5 nal_unit setup_unit(setup_unit_length); } } // additional fields } [0065] In some examples, a track uses the below VolumetricVisualSampleEntry with a sample entry type of'v3e1' or'v3eg', e.g., as described in ISO/IEC 23090-10. This entry contains a V3CConfigurationBox (912), and optional boxes for 2D video decoder configuration information, such as HEVCconfiguartionBox,VVCconfiguartionBox, as defined in ISO/IEC 14496-15, may be present to signal the information for the 2D video decoder configuration and initialization. Sample Entry Type: 'v3e1', 'v3eg' Docket No. D23136WO01 Container: SampleDescriptionBox Mandatory: A 'v3e1' or 'v3eg' sample entry is mandatory Quantity: One or more aligned(8) class V3CBitstreamSampleEntry() extends VolumetricVisualSampleEntry (type) { // type is 'v3e1' or 'v3eg' V3CConfigurationBox config; Box[] any_box; // optional } [0066] FIGS.10A-10B are block diagrams illustrating multiple track encapsulation of the V3C bitstream (672) according to some examples. In the examples shown, multiple track encapsulation is implemented using two types of tracks, i.e., a V3C atlas track (1002) and a V3C video component track (1004). The V3C atlas track (1002) contains a V3C parameter set and may also contain atlas parameter sets in the sample entry and atlas component bitstream NAL units in the samples. The V3C video component track (1004) contains access units of video- coded elementary streams for packed texture and transparency data in the samples. Thus, one or two V3C atlas tracks (1002) and at least one packed video track may be present in a video file. In the example shown in FIG.10A, the atlas sub-stream is carried in two V3C atlas tracks (10021, 10022), one being used for common atlas data and the other being used for individual atlas data. In the example shown in FIG.10B, the entire atlas sub-stream is carried into a single track (1002). [0067] In some examples, the V3C atlas tracks (1002) use V3CAtlasSampleEntry, which extends VolumetricVisualSampleEntry with a sample entry type of'v3c1', 'v3cg','v3cb','v3a1', or 'v3ag'. A V3C atlas track sample entry contains a V3CConfigurationBox (912), and an optionalV3CUnitHeaderBox contains the V3C unit header describing the data carried by the corresponding track. Sample Entry Type: 'v3c1', 'v3cg', 'v3cb', 'v3a1', or 'v3ag' Container: SampleDescriptionBox Mandatory: A 'v3c1', 'v3cg', 'v3cb', 'v3a1', or 'v3ag' sample entry is mandatory Quantity: One or more aligned(8) class V3CAtlasSampleEntry() extends VolumetricVisualSampleEntry (type) { // type is 'v3c1', 'v3cg', 'v3cb', 'v3a1', or 'v3ag' Docket No. D23136WO01 V3CConfigurationBox config; V3CUnitHeaderBox unit_header;//optional } [0068] In some examples, a V3CUnitHeaderBox is also present in scheme_specific_data Box array of the SchemeInformationBox of V3C packed video tracks. The V3CUnitHeaderBox contains the V3C unit header describing the data carried by the video track. aligned(8) class V3CUnitHeaderBox extends FullBox('vunt', version = 0, 0){ v3c_unit_header header(); } header contains a single instance of the 32-bit V3C unit header. [0069] In some examples, to link a track to the other tracks, the track reference tool defined in ISO/IEC 14496-12 may be used. To link a V3C atlas tack carrying common atlas data only to the V3C atlas tracks carrying individual atlas data, the 4CCs of these track reference types is 'v3cs'. To link a V3C atlas tack carrying individual atlas data to the V3C packed video track, the 4CCs of these track reference types is 'v3cp'. [0070] FIG.11 is a block diagram illustrating further details of the multiple track encapsulation corresponding to the example illustrated in FIG.10A. The atlas sub-stream (620) is carried into the atlas tracks (10021, 10021). The atlas track (10021) is used for common atlas data, and the atlas track (10022) is used for individual atlas data. The video sub-stream (662) is carried in the separate track (1004). As indicated in FIG.11, each of the atlas tracks (10021, 10021) contains the boxes V3CConfigurationBox andV3CUnitHeaderBox, with the latter having a 32-bit unit header of the V3C unit which represents the samples in the track. The V3C packed video track(1004) contains video configuration information for the 2D video decoder’s configuration and initialization, and the boxV3CUnitHeaderBox which contains the unit header of the V3C unit representing samples in the track. [0071] FIG.12 is a block diagram illustrating further details of the multiple track encapsulation corresponding to the example illustrated in FIG.10B. In this example, the entire atlas sub-stream (620) is carried in the single track (1002). As indicated in FIG.12, each sample in the V3C atlas track (1002) corresponds to one or more ACL NAL units, and each sub-sample contains one ACL NAL unit. The V3C atlas track (1002) contains one box V3CConfigurationBox, which includes the V3C bitstream’s decoding specific information, Docket No. D23136WO01 and one boxSubSampleInformationBox, which lists the track sub-samples. The 32-bit unit header of the V3C unit, which represents the sub-sample, is copied to the codec_specific_parameters field of the sub-sample entry in the box SubSampleInformationBox. The V3C unit type of each sub-sample is identified by parsing the codec_specific_parameters field of the sub-sample entry in the box SubSampleInformationBox. [0072] The following portion of this specification describes storage and metadata signaling of a 2D video bitstream carrying MPI video. In particular, we describe encapsulation in a track of the MPI bitstream, which is coded by a suitable 2D video codec as described above. The relevant box types are listed in bold in Table 2. Table 2: Box types described in this section and their relation to boxes not specified in this document. [0073] In some examples, the storage of a video bitstream (MPI bitstream) into a track uses some of the existing capabilities of the ISO base media file format, e.g., as specified in ISO/IEC 14496-12 and 14496-15, but also defines an extension thereof to support the following features of the MPI video. - Indication that decoded pictures are packed pictures containing texture and transparency maps of MPI layers corresponding to an MPI representation; - The packing information of MPI-coded video frames, such as a representation of two spatially packed constituent frames or two temporally interleaved constituent frames. [0074] In various examples, a player that does not recognize the scheme type in the box SchemeTypeBox may be configured to ignore the corresponding track. A player, e.g., the Docket No. D23136WO01 player (590), that recognizes the scheme type in the boxSchemeTypeBox operates to parse all boxes contained in the box SchemeInformationBox to determine whether it has capabilities needed to properly process the track. A player is configured to ignore the track unless it supports all boxes contained in the boxSchemeInformationBox and all syntax elements and syntax element values present in those boxes. [0075] The use of the MPI video scheme for the restricted video sample entry type 'resv' indicates that the decoded pictures are packed pictures containing MPI video content. The use of the MPI video scheme is indicated byscheme_type equal to 'mpiv' within SchemeTypeBox in the boxRestrictedSchemeInfoBox. Box Type: 'mpiv' Container: SchemeInformationBox Mandatory: Yes, when scheme_type is equal to 'mpiv' Quantity: Zero or one [0076] The boxMPIVideoBox is used to indicate information specific to the MPI video as described above. In various examples, this box is a container box that contains boxes indicating information for the following: (1) The packing information of the decoded picture that the decoded frames either contain a representation of two spatially packed constituent frames or contain a representation of two temporally interleaved constituent frames which form an MPI representation. (2) Region-wise packing information, when applicable. [0077] The boxMPIVideoBox contains the box MPIPackingFormatBox, which specifies the arrangement type of texture and transparency maps of the MPI layers in the decoded pictures and whether the arrangement is applied per layer or picture among the constituent pictures. The box MPIVideoBox optionally contains the box MPIRegionWisePackingBox for mapping between the packed regions and the corresponding MPI layers. Pertinent definitions are as follows: aligned(8) class MPIVideoBox extends Box('mpiv') { MPIPackingFormatBox() packing_format_box; // mandatory MPIRegionWisePackingBox() region_packing; // optional // optional boxes but no fields } Docket No. D23136WO01 aligned(8) class MPIPackingFormatBox() extends FullBox('mppf', 0, 0) { MPIPackingFormatStruct() packing_format_struct; } aligned(8) class MPIPackingFormatStruct(){ bit(5) reserved = 0; unsigned int(2) packing_type; unsigned int(1) packing_unit; } Herein, packing_type indicates the type of the arrangement of texture and transparency map of MPI layers in decoded pictures. The following values are specified and one of these values are to be set: Table 3: Packing types When the value ofpacking_type indicates the temporal interleaving frame packing arrangement, the player needs to implicitly set the composition timestamp for constituent picture 0 to coincide with the composition timestamp for constituent picture 1. packing_unit indicates whether the arrangement type indicated by packing_type is applied per picture among constituent pictures or per layer.packing_unit equal to 0 indicates that the arrangement type indicated by packing_type is applied per picture among constituent pictures. It should be equal to 0 when packing_type is equal to 1. Docket No. D23136WO01 packing_unit equal to 1 indicates that the arrangement type indicated by packing_type is applied per layer. NOTE 1: FIG.8A illustrates an example of packing_type equal to 3 and packing_unit equal to 0. FIG.8D illustrates an example of packing_type equal to 2 andpacking_unit equal to 0. NOTE 2: FIG.8B illustrates an example ofpacking_type equal to 3 and packing_unit equal to 1. FIG.8C illustrates an example of packing_type equal to 2 and packing_unit equal to 0. [0078] The box MPIRegionWisePackingBox specifies the mapping between packed regions and the corresponding layers and specifies the location and size of each region. Pertinent definitions are as follows: Box Type: 'mprw' Container: MPIVideoBox Mandatory: No Quantity: Zero or one aligned(8) class MPIRegionWisePackingBox extends FullBox('mprw', 0, 0) { MPIRegionWisePackingStruct() region_wise_packing_struct; } aligned(8) class MPIRegionWisePackingStruct() { unsigned int(8) num_mpi_layers; unsigned int(16) mpi_layer_width; unsigned int(16) mpi_layer_height; for (i = 0; i < 2*num_mpi_layers; i++) { unsigned int(7) mpi_layer_id[i]; unsigned int(1) mpi_pic_type[i]; MPIRegionPacking mpi_packed_region[i]; } } num_mpi_layers specifies the number of MPI layers corresponding to the signaled packed regions in this syntax structure. Docket No. D23136WO01 mpi_layer_width andmpi_layer_height specify the width and height, respectively, of the MPI layer region within the original map picture before packing. mpi_layer_id[i] specifies that the identifier of MPI layers corresponding to the i-th packed region. mpi_pic_type[i] indicates whether the i-th packed region corresponds to transparency or texture. The 0 value ofmpi_pic_type[i] specifies that the i-th packed region corresponds to the texture of the MPI layer indicated bympi_layer_id[i]. The 1 value of mpi_pic_type[i] specifies that the i-th packed region corresponds to transparency of the MPI layer indicated bympi_layer_id[i]. mpi_packed_region[i] specifies the region-wise packing between the i-th packed region and the MPI layer indicated by mpi_layer_id[i]. aligned(8) class MPIRegionPacking { unsigned int(16) mpi_reg_width; unsigned int(16) mpi_reg_height; unsigned int(16) mpi_reg_top; unsigned int(16) mpi_reg_left; unsigned int(16) packed_reg_width; unsigned int(16) packed_reg_height; unsigned int(16) packed_reg_top; unsigned int(16) packed_reg_left; } [0079] TheMPIRegionPacking structure specifies an MPI layer region and a respective packed region. The parameters mpi_reg_width,mpi_reg_height, mpi_reg_top, and mpi_reg_left specify the width, height, top offset, and left offset, respectively, of the MPI layer region within the original map picture before packing. The parameters packed_reg_width,packed_reg_height,packed_reg_top, and packed_reg_left specify the width, height, the offset, and the left offset, respectively, of the packed region, either within spatially packed texture and transparency constituent pictures or within each constituent picture of temporally interleaved texture and transparency constituent pictures. [0080] FIG.13 is a block diagram illustrating single-track encapsulation of the 2D video stream (MPI bitstream) according to some examples. As indicated in FIG.13, the MPI bitstream Docket No. D23136WO01 (508) is carried to a single track (1302). The track (1302) contains 2D video configuration information for the 2D video decoder configuration and initialization and also contains the box MPIVideoBox, which has the packing information of texture and transparency maps of the MPI layers in decoded pictures and mapping between the packed regions and the respective MPI layers. The access units of video-coded elementary streams for packed texture and transparency data are stored in the samples. MPI Video Encapsulation and Signaling Using DASH [0081] An important component at the server side is the MPD. In various examples, the DASH MPD generator includes MPI video-specific descriptors. Those descriptors include the metadata described above, e.g., specifying how texture and transparency maps of MPI layers are packed to decoded frames. When the user joins the session to start the MPI playback, an MPD will be delivered to the player (590)/client. By parsing the metadata from the MPD, e.g., for packing information, the player (590) determines which Adaptation Set and Representation contain the supported MPI video bitstream and which Adaptation Set and Representation cover the current viewing orientation at the highest quality and at a bitrate that may be afforded by the prevailing estimated network throughput. The player (590) then issues segment requests accordingly. [0082] The single-track mode in DASH enables streaming of segments where the V3C bitstream is stored using single-track encapsulation, as described above. [0083] The single-track mode in DASH can be represented as one Adaptation Set with one or more Representations. Each representation is associated with one particular encoded MPI bit rate. The Initialization Segment of representations contains 2D video decoder information for configuration and initialization of the 2D video decoder and all parameter sets needed to initialize the V3C decoder, including the V3C parameter sets as well as other parameter sets for sub-bitstreams. Media segments of the representation contain one or more track fragments of the V3C bitstream track, e.g., as described in the preceding section, and sub-sample information, which lists the sub-samples to support sub-sample level access in the segment. [0084] An example of MPEG DASH signaling when an MPI video representation is encoded in the single-track mode with HEVC is provided below: <?xmI versIon=”1.0” encoding="UTF-8”?> <MPD xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns="urn:mpeg:dash:schema:mpd:2011" Docket No. D23136WO01 xsi :schemaLocation="urn:mpeg:dash:schema:mpd:2011 DASH-MPD.xsd" type="static" mediaPresentationDuration="PT3256S" minBufferTime="PT1.2S" profiles="urn:mpeg:dash:profile:isoff-on-demand:2011"> <BaseURL>http://cdn1.example.com/</BaseURL> <Period> <!- MPI AdaptationSet --> <AdaptationSet id="mpi" mimeType="video/mp4" codecs="v3e1.L2.0.0.1, resv.vvvc.hvc1" frameRate="30"> <SegmentList> <Initialization sourceURL="seg-m-init.mp4"/> </SegmentList> <Representation bandwidth="512000"> <BaseURL>mpi-512k.mp4</BaseURL> </Representation> <Representation bandwidth="1024000"> <BaseURL>mpi-1024k.mp4</BaseURL> </Representation> <Representation bandwidth="2048000"> <BaseURL>mpi-2048k.mp4</BaseURL> </Representation> </ AdaptationSet> </Period> </MPD> [0085] In the multi-track mode, the common atlas, individual atlas, and packed video component are represented in the DASH manifest (MPD) file as a separate Adaptation Set. An Adaptation Set for common atlas information serves as the Main Adaptation Set. In some examples, the Main Adaptation Set contains a single Initialization Segment at the adaptation set level. The Initialization Segment contains all parameter sets needed to initialize the V3C decoder, including the V3C parameter sets, as well as other parameter sets for the component sub-bitstreams. Media Segments for the Representation of the Main Adaptation Set contain one Docket No. D23136WO01 or more track fragments of the V3C atlas track. Media Segments for the Representations of Video Component Adaptation Sets contains one or more track fragments of the corresponding packed video track as described above. [0086] FIG.14 is a block diagram illustrates a DASH configuration for grouping atlas and packed video components belonging to the same MPI content within an MPEG-DASH MPD file according to some examples. As indicated in FIG.14, a preselection (1402) is used to combine multiple Adaptation Sets (e.g., 1404, 1406, 1408) into a single decoding instance and MPI user experience. In some examples, the preselection (1402) may be signaled in the MPD, e.g., using aPreSelection element within the Period element or a Preselection descriptor. Each of the common atlas track, individual atlas track, and packed video track are presented as a separate respective Adaptation Set. The preselection (1402) indicates the grouping of three Adaptation Sets (1404, 1406, 1408) to provide the user with a single MPI experience. [0087] To identify the type of Video Component Adaptation Set, a V3CVideoComponent descriptor is used. A V3CVideoComponent descriptor is an EssentialProperty element with the@schemeIdUri set to "urn:mpeg:mpegI:v3c:2020:videoComponent". [0088] At the Adaptation Set level, the V3CVideoComponent descriptor is signaled for each V3C video component that is present in the Representations of the Video Component Adaptation Set. The @value of the V3CVideoComponent descriptor may not be present. In some examples, the V3CVideoComponent descriptor includes elements and attributes specified in Table 4. Table 4: Elements and attributes for the V3CVideoComponent descriptor. Docket No. D23136WO01 Docket No. D23136WO01 Docket No. D23136WO01 [0089] An example XML schema for the V3CVideoComponent descriptor is shown below: <?xml version="1.0" encoding="UTF-8"?> <xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema" targetNamespace="urn:mpeg:mpegI:v3c:2020" xmlns:v3c="urn:mpeg:mpegI:v3c:2020" elementFormDefault="qualified"> <xs:element name="videoComponent" type="v3c:VideoComponentType"/> <xs:complexType name="VideoComponentType"> <xs:attribute name="type" type="xs:string" use="required" /> <xs:attribute name="is_auxiliary" type="xs:boolean" default=”false”/> <xs:attribute name="map_index" type="xs:integer" /> <xs:attribute name="attribute_type" type="xs:unsignedByte" /> <xs:attribute name="attribute_index" type="xs:unsignedByte" /> <xs:attribute name="attribute_dim_partition_index" type="xs:unsignedByte" /> <xs:attribute name="atlas_id" type="xs:integer" use="optional" default="0" /> <xs:attribute name="tile_ids" type="UIntVectorType" use="optional" /> </xs:complexType> </xs:schema> [0090] At the Adaptation Set level of packed video, one V3CVideoComponent descriptor is signaled with videoComponent@type as “packed”. Alternatively, to indicate which attribute components are packed into the Adaptation Set, two V3CVideoComponent descriptors are signaled: one withvideoComponent@type as “packed” and videoComponent@attribute_type as “ATTR_TEXTURE”, and the other with videoComponent@type as “packed” and videoComponent@attribute_type as “ATTR_TRANSPARENCY”. [0091] The following MPD example illustrates the signaling component of each Adaptation Set, such as the packed video Adaptation Set, and how the preselection element (1402) is used to indicate the group of Adaptation Sets that represents a single MPI video. Docket No. D23136WO01 <?xml version="1.0" encoding="UTF-8”?> <MPD xmlns="urn:mpeg:dash:schema:mpd:2011" xmlns:v3c="urn:mpeg:mpegl:v3c:2020" type="static" mediaPresentationDuration="PT10S" minBufferTime="PT1S" profiles="urn:mpeg:dash:profile:isoff-on-demand :2011"> <Period> <!-- Main V3C AdaptationSet --> <AdaptationSet mimeType="catlas/mp4" id="1" codecs="v3cb"> <EssentialProperty schemeldUri ="urn:mpeg:dash:preselection :2016" /> <Representation> … </Representation> </ AdaptationSet> <!- Atlas AdaptationSet--> <AdaptationSet id ="2"mimeType="atlas/mp4" codecs="v3a1" dependencyld="1" > <EssentialProperty schemeldUri ="urn:mpeg:dash:preselection :2016" /> <Representation> … </Representation> </Adaptation Set> <!-Packed video AdaptationSet--> <AdaptationSet id="3" mimeType="video/mp4" codecs="resv.vvvc.hvc1"> <EssentialProperty schemeldUri ="urn:mpeg:dash:preselection :2016" /> <Essentia!Property schemeldUri ="urn:mpeg:mpegl:v3c:2020:component"> <v3c:videoComponent type="packed" /> </Essential Property> <Representation> … </Representation> </ AdaptationSet> Docket No. D23136WO01 <!-- Preselections --> <Preselection id="1" tag ="1" preselectionComponents="123" codecs="v3cb"> <!-V3C Descriptor--> <EssentialProperty schemeldUri="urn:mpeg:mpegl:v3c:2020:vpc" vld ="1" /> </Preselection> </Period> </MPD> [0092] The next portion of this section describes streaming (518) of segments when the corresponding MPI bitstream (508) is stored using the track encapsulation described above. In some examples, each MPI representation is associated with one particular encoded MPI video having a particular bitrate. The Initialization Segment of the representations contains 2D video decoding-specific information for configuration and initialization of the 2D video decoder and the box MPIVideoBox described herein. Media Segments for the Representations contain one or more track fragments of the track as described above. [0093] The DASH MPD generator includes MPI video-specific descriptors. The descriptors include the packing type, region-wise packing information. This information may be generated based on equivalent information in the segments. [0094] To identify the packing arrangement of texture and transparency maps of the MPI Representation at the Adaptation Set, an MPIPackingFormat (MPF) descriptor is used. The MPF descriptor is an element the EssentialProperty descriptor with the atribute @schemeIdUri being set to "urn:mpeg:dash:mpi:2023:pf." In some examples, one MPF descriptor is present at the adaptation set level. One MPF descriptor may also be present at the representation level. [0095] In some examples, the @value attribute of the MPF descriptor is not present. The MPF descriptor may include elements and attributes specified in Table 5. Table 5. Elements and attributes for the MPIPackingFormat (MPF) descriptor. Docket No. D23136WO01 [0096] In some examples, an XML schema for the MPF descriptor can be as follows: <?xml version="1.0" encoding="UTF-8"?> elementFormDefault="qualified"> <xs:element name="mpiPackingFormat" type="mpiPackingFormatType"/> <xs:complexType name="mpiPackingFormatType "> <xs:attribute name="packing_type" type="xs:unsignedByte" use="required" /> </xs:complexType> </xs:schema> [0097] The following MPD example illustrates the MPD signaling when the decoded output pictures of each representation represent texture and transparency constituent pictures in a top-to-bottom packing arrangement. <?xml version="1.0" encoding= “UTF-8”?> <MPD xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" Docket No. D23136WO01 xmlns="urn:mpeg:dash:schema:mpd:2011" xsi :schemaLocation="urn:mpeg:dash:schema:mpd:2011 DASH-MPD.xsd" type="static" mediaPresentationDuration=" PT3256S" minBufferTime="PT1.2S" profiles="urn:mpeg:dash:profile :isoff-on-demand:2011"> <BaseURL>http://cdnl.example.com/</BaseURL> <Period> <!- MPI AdaptationSet --> <AdaptationSet id="mpi" mimeType="video/mp4" codecs="resv.mpiv. vvc1" frameRate="30"> <EssentialProperty schemeldUri ="urn:mpeg:dash:mpi:2023:pf"> <mpiPackingFormat packing_type="2" /> <Segmentlist> <InitiaIization sourceURL="seg-m-init.mp4"/> </Segmentlist> <Representation bandwidth="512000"> < BaseURL> mpi-512k.mp4</BaseURL> </Representation> <Representation bandwidth="1024000"> < BaseURL> mpi-1024k.mp4</BaseURL> </Representation> <Representation bandwidth="2048000"> < BaseURL> mpi-2048k.mp4</BaseURL> </Representation> </AdaptationSet> </Period> </MPD> [0098] The multi-view MPI experience can be achieved by switching reference views as described above in reference to FIG.4. The 3D scene has a total of N reference camera views. In the example illustrated in FIG.4, N=40. To render a novel view at time instance t0, only the four neighboring reference views within the block (410) are used. The novel view position moves along the dotted path (402) over time. Along the path (402) at time instance tn, the novel view’s nearest neighbors have changed to be the reference views within the block (420). Docket No. D23136WO01 [0099] In the multiple view MPI system, we assume that every single view MPI is independently coded and encapsulated in tracks or segments, e.g., as described above. Since the user may request a different novel viewport along the timeline, the player (590) needs to be configured to dynamically switch between different MPI bitstreams. Thus, among the multiple view MPIs, the player (590) is configured to identify which MPI views to be prefetched and to switch to a proper MPI views accordingly when the viewport is changes. For this switching operation, the information to be signaled in the MPD can be as follows. • The entire MPI structure including the number of MPI views. • The information corresponding to each reference camera view. o Reference camera view identifier; o Extrinsic camera and intrinsic camera information; o Vertical and horizontal field of view information of a reference camera view. • Neighboring reference views for joint rendering. [0100] In some examples, we map each level of our multi-view MPI streaming system to a hierarchical data structure in the MPD. In some examples, this structure can be as follows. • InitializationSet: The InitializationSet contains the URL of an initialization segment which contains all extrinsic and intrinsic camera information and the needed post-decoder operation-specific information, such as packing information, per time period. Note that different MPI scenes might have different numbers of cameras and different camera poses set up as time goes by. In such cases, multiple InitalizationSet attribures are present, and each InitializationSet attribute contains a respective URL of the initialization segment containing extrinsic and intrinsic camera information for the time period. Each adaptation set has a link to the corresponding InitializationSet for the time period. • Preselection: Each preselection represents a combination of Adaptation Sets that can be selected for joint rendering in the given time period. If we have 20 camera views, with up to four neighboring views being able to improve the novel viewport rendering, the pre- selections indicate the corresponding groups of neighboring reference views for joint rendering. When the pre-selections are present in the MPD, the player (590) can seamlessly switch to relevant views by identifying which combination of adaptation sets needs to be prefetched and requesting the relevant multiple adaptation sets according to the viewport change. • Adaptation Set: Each adaptation set represents one MPI from one reference camera view. Each reference camera view is associated with extrinsic and intrinsic camera information. If we have 20 camera views, then we will have 20 corresponding adaptation sets. Docket No. D23136WO01 • Representation: Each representation is associated with one particular bitrate version of the encoded MPI. Different respective representation is chosen according to the current network condition existing between the server and the client. • Segment: There are two types of segments: initialization segments and media segments. These may be encapsulated as already described above. [0101] FIG.15 is a flowchart (1500) illustrating communications between a server (1520) and a client (1530) and the corresponding processing operations performed at the client (1530) according to various examples. In some examples, the client (1530) may be or include the player (590). A final block (1511) of operations in the flowchart (1500) includes the client generating the output view and rendering the generated view according to the user viewport. [0102] When the user starts or joins the server-client session to start MPI playback, an MPD (1501) is sent from the server (1520) to the client (1530). The client (1530) parses MPD (1501) in a first block (1502) of operations and determines a preselection based on the present viewport in a second block (1503) of operations. In a third block (1504) of operations, the client (1530) selects a collection of adaptation sets, sends a request (1505) for the corresponding initialization segment, and receives the requested initialization segment (1506). In a fourth block (1507) of operations, the client (1530) determines one or more representations and sends one or more corresponding requests (1508) to the server (1520). In a fifth block (1510) of operations, the client (1530) receives the requested one or more media segments (1509) and decodes the corresponding bitstreams. In the final block (1511), the client (1530) uses the decoded bitstreams to generate the view(s) that fit(s) the current view pose and renders the generated view according to the user viewport. [0103] The InitializationSet contains the URL of the initialization segment (e.g., of 1506) which contains extrinsic and intrinsic camera information of all MPI views and specifies the needed MPI post-decoder operation-specific information, such as the packing information, for the time period. When that initialization information is updated as time goes by, multiple instances of the InitalizationSet (1506) may be present, and each such instance of InitializationSet contains the URL of the corresponding initialization segment containing the updated initialization information including extrinsic and intrinsic camera information and MPI post-decoder operation-specific information for the time period. Each adaptation set has @InitializationSetRef to be associated with a proper InitializationSet for the time period. [0104] Preselection for a combination of Adaptation Sets that can be selected for joint rendering for the time period may either be signaled in the MPD (1501) using a Docket No. D23136WO01 PreSelection element within thePeriod element or in the Preselection descriptor at the Adaptation Set level. A PreSelection element is signaled with an ID list for the @preselectionComponents attribute including the ID of the Adaptation Set for the neighboring reference MPI views. ThePreSelection element is associated with a 3D spatial region coverage (SRC) descriptor to indicate the 3D spatial region covered by the combination of MPI views in the Preselection. [0105] A SupplementalProperty element with an @schemeIdUri attribute equal to "urn:mpeg:dash:mpi:2023:src" is referred to as the 3D spatial region coverage (SRC) descriptor. In various examples, one SRC descriptor is present at the Preselection level. An SRC descriptor may not be present at the MPD, adaptation set, or representation level. The SRC descriptor indicates the 3D spatial region covered by the combination of MPI views in the Preselection. [0106] In some examples, the @value attribute of the SRC descriptor may not be present. In various examples, the SRC descriptor may include some or all of the elements and attributes specified in Table 6. Table 6. Elements and attributes for the SRC descriptor Docket No. D23136WO01 Docket No. D23136WO01 [0107] In some examples, an XML schema for the SRC descriptor is as follows: <?xml version="1.0" encoding="UTF-8"?> <xs:schema xmlns:xs=http://www.w3.org/2001/XMLSchema elementFormDefault="qualified"> <xs:complexType name="spatialRegionType"> <xs:attribute name="type" type="xs:unsignedByte" use="optional" default="0" <xs:element name="cuboid" type="spatialRegionCuboidType" minOccurs="0" maxOccurs="1"/> <xs:element name="viewport" type="ViewportType" minOccurs="0" maxOccurs="1"/> </xs:complexType> <xs:complexType name="spatialRegionCuboidType"> <xs:attribute name="anchor" type="UIntVectorType" use="required" minLength="3" maxLength="3" /> <xs:attribute name="dimensions" type="UIntVectorType" use="required" minLength="3" maxLength="3"/> </xs:complexType> <xs:complexType name="ViewportInfoType"> <xs:attribute name="vp_pos" type="FloatVectorType" use="required" minLength="3" maxLength="3"/> <xs:attribute name="vp_quat" type="IntVectorType" use="required" minLength="3" maxLength="3"/> </xs:complexType> </xs:schema> Docket No. D23136WO01 [0108] In some examples, to identify the MPI view information employed, an MPIView descriptor is used. An MPIView descriptor is an EssentialProperty or SupplementalProperty element with an@schemeIdUri attribute set to "urn:mpeg:dash:mpi:2023:view". At most one MPIView descriptor is present at the adaptation set level. [0109] In some examples, the @value attribute of the MPIView descriptor may not be present. In various examples, the MPIView descriptor includes some or all of the elements and attributes specified in Table 7. Table 7: Elements and attributes for the MPIView descriptor. Docket No. D23136WO01 Docket No. D23136WO01 [0110] In some examples, an XML schema for the MPIView descriptor is as follows. <?xml version="1.0" encoding="UTF-8"?> <xs:schema xmlns:xs=http://www.w3.org/2001/XMLSchema elementFormDefault="qualified"> <xs:attribute name="view_id" type="xs:integer" use="required" /> <xs:element name="ViewportInfo" type="v3c:ViewportInfoType"/> <xs:element name="mpiPackingFormat" type="mpiPackingFormatType"/> <xs:complexType name="ViewportInfoType"> <xs:attribute name="vp_pos" type="FloatVectorType" use="required" minLength="3" maxLength="3"/> <xs:attribute name="vp_quat" type="IntVectorType" use="required" minLength="3" maxLength="3"/> <xs:attribute name="cam_type" type="xs:unsignedByte" use="optional”/> <xs:element name="erp" type="ViewInfoERPType" minOccurs="0" maxOccurs="1"/> Docket No. D23136WO01 <xs:element name="per" type="ViewInfoPerspectiveType" minOccurs="0" maxOccurs="1"/> <xs:element name="ortho" type="ViewInfoOrthoType " minOccurs="0" maxOccurs="1"/> </xs:complexType> <xs:complexType name=" ViewInfoERPType"> <xs:attribute name="phi" type="FloatVectorType" use="required" minLength="2" maxLength="2" /> <xs:attribute name="theta" type="FloatVectorType" use="required" minLength="2" maxLength="2" /> </xs:complexType> <xs:complexType name="ViewInfoPerspectiveType"> <xs:attribute name="focal" type="FloatVectorType" use="required" minLength="2" maxLength="2" /> <xs:attribute name="principal" type="FloatVectorType" use="required" minLength="2" maxLength="2" /> </xs:complexType> <xs:complexType name="ViewInfoOrthoType"> <xs:attribute name="ortho" type="FloatVectorType" use="required" minLength="2" maxLength="2" /> </xs:complexType> </xs:schema> [0111] FIG.16 is a block diagram illustrating a DASH configuration that can be used in the MPI video system (500) for the support of seamless switching among MPI views within an MPEG-DASH MPD file according to some examples. In the example shown, the MPD (1501) includes an InitializationSet (1601) and a Period (1602). The InitializationSet (1601) includes the URL of the initialization segment containing all needed initialization information. The Period (1602) includes one or more Preselections (1603k) and one or more Adaptation Sets (1604n). A Preselection (1603k) indicates the combination of adaptation sets corresponding to the neighboring reference view(s) that can be used to improve the novel view generation by joint rendering. The 3D spatial region coverage information by such combination of adaptation sets is indicated with the spatial region coverage descriptor. Each Adaptation Set (1604n) is also Docket No. D23136WO01 indicated with the MPI view information including the corresponded extrinsic and intrinsic camera information. Example Hardware [0112] FIG.17 is a block diagram illustrating an example computing device (1700) according to various examples. In some examples, two or more instances of the computing device (1700) are used in the MPI video system (500). [0113] The computing device (1700) of FIG.17 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device (1700) may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices (1702) and one or more storage devices (1704)). Additionally, in various embodiments, the computing device (1700) may not include one or more of the components illustrated in FIG.17, but may include interface circuitry for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other appropriate interface). For example, the computing device (1700) may not include a display device (1710), but may include display device interface circuitry (e.g., a connector and driver circuitry) to which an external display device (1710) may be coupled. [0114] The computing device (1700) includes a processing device (1702) (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device” interchangeably refer to any device or portion of a device that processes electronic data from registers and/or memory to transform that electronic data into other electronic data that may be stored in registers and/or memory. In various embodiments, the processing device 1702 may include one or more digital signal processors (DSPs), application- specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices. [0115] The computing device (1700) also includes a storage device (1704) (e.g., one or more storage devices). In various embodiments, the storage device (1704) may include one or more memory devices, such as random-access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based Docket No. D23136WO01 memory devices, solid-state memory devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device (1704) may include memory that shares a die with the processing device (1702). In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory (STT-MRAM), for example. In some embodiments, the storage device (1704) may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g., the processing device (1702)), cause the computing device (1700) to perform any appropriate ones of the methods disclosed herein below or portions of such methods. [0116] The computing device (1700) further includes an interface device (1706) (e.g., one or more interface devices (1706)). In various embodiments, the interface device (1706) may include one or more communication chips, connectors, and/or other hardware and software to govern communications between the computing device (1700) and other computing devices. For example, the interface device (1706) may include circuitry for managing wireless communications for the transfer of data to and from the computing device (1700). The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device (1706) for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and/or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device (1706) for managing wireless communications may operate in accordance with a Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. In some embodiments, circuitry included in the interface device (1706) for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E- UTRAN). In some embodiments, circuitry included in the interface device (1706) for managing wireless communications may operate in accordance with Code Division Multiple Access Docket No. D23136WO01 (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device (1706) may include one or more antennas (e.g., one or more antenna arrays) configured to receive and/or transmit wireless signals. [0117] In some embodiments, the interface device (1706) may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device (1706) may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device (1706) may support both wireless and wired communication, and/or may support multiple wired communication protocols and/or multiple wireless communication protocols. For example, a first set of circuitry of the interface device (1706) may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry of the interface device (1706) may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry of the interface device (1706) may be dedicated to wireless communications, and a second set of circuitry of the interface device (1706) may be dedicated to wired communications. [0118] The computing device (1700) also includes battery/power circuitry (1708). In various embodiments, the battery/power circuitry (1708) may include one or more energy storage devices (e.g., batteries or capacitors) and/or circuitry for coupling components of the computing device (1700) to an energy source separate from the computing device (1700) (e.g., to AC line power). [0119] The computing device (1700) also includes a display device (1710) (e.g., one or multiple individual display devices). In various embodiments, the display device (1710) may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display. [0120] The computing device (1700) also includes additional input/output (I/O) devices (1712). In various embodiments, the I/O devices (1712) may include one or more data/signal transfer interfaces, audio I/O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration Docket No. D23136WO01 sensors, etc.), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc. [0121] Depending on the specific embodiment, various components of the interface devices (1706) and/or I/O devices (1712) can be configured to output suitable control signals, receive suitable control/telemetry signals, and receive and transmit data streams. In some examples, the interface devices (1706) and/or I/O devices (1712) include one or more analog-to- digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device (1702) and/or the storage device (1704). In some additional examples, the interface devices (1706) and/or I/O devices (1712) include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device (1702) and/or the storage device (1704) into an analog form suitable for being transmitted through a communication channel. [0122] According to an example embodiment disclosed above, e.g., in the summary section and/or in reference to any one or any combination of some or all of FIGs.1-17, provided is an apparatus for streaming an MPI video, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: generate a sequence of video frames, each of the video frames including a respective plurality of patches representing texture and transparency layers of one or more multiplane images of the MPI video; apply video compression to the sequence of video frames to generate a video sub- stream; generate a sequence of representations of atlas frames corresponding to the sequence of video frames to specify at least a packing arrangement of the patches; apply compression to the sequence of representations to generate an atlas sub-stream; and multiplex the video sub-stream and the atlas sub-stream to generate a first coded bitstream encoding at least a portion of the MPI video. [0123] According to another example embodiment disclosed above, e.g., in the summary section and/or in reference to any one or any combination of some or all of FIGs.1-17, provided is a method for streaming an MPI video, the method comprising: generating a sequence of video frames, each of the video frames including a respective plurality of patches representing texture and transparency layers of one or more multiplane images of the MPI video; applying video compression to the sequence of video frames to generate a video sub-stream; generating a sequence of representations of atlas frames corresponding to the sequence of video frames to specify at least a packing arrangement of the patches; applying compression to the sequence of Docket No. D23136WO01 representations to generate an atlas sub-stream; and multiplexing the video sub-stream and the atlas sub-stream to generate a first coded bitstream encoding at least a portion of the MPI video. [0124] In some embodiments of the above method, the first coded bitstream is a visual volumetric video-based coding (V3C) bitstream in accordance with the ISO.IEC.23090-5 standard. [0125] In some embodiments of any of the above methods, the method further comprises encapsulating the first coded bitstream into a media file configured for a file playback. [0126] In some embodiments of any of the above methods, the method further comprises: encapsulating the first coded bitstream into a sequence of segments according to a selected media container file format; and streaming the sequence of segments over a communication channel to a client device for playback. [0127] In some embodiments of any of the above methods, the first coded bitstream is configured for encapsulation in a single V3C bitstream track. [0128] In some embodiments of any of the above methods, the first coded bitstream is configured for encapsulation in two or more V3C bitstream tracks. [0129] In some embodiments of any of the above methods, the two or more V3C bitstream tracks include: a first track that carries at least a portion of the atlas sub-stream; and a second track that carries at least a portion of the video sub-stream. [0130] In some embodiments of any of the above methods, texture and transparency patches of the respective plurality of patches are packed into a corresponding video frame using a packing arrangement selected from the group consisting of: a side-by-side arrangement; a top- to-bottom arrangement; a vertically interleaved arrangement; and a horizontally interleaved arrangement. [0131] In some embodiments of any of the above methods, the texture and transparency patches include patches of a first size and patches of a different second size. [0132] In some embodiments of any of the above methods, the respective plurality of patches has fewer than all of the patches of a corresponding multiplane image. (Partial packing.) [0133] In some embodiments of any of the above methods, the method further comprises: providing to a client device a media presentation description (MPD) of an MPI streaming content stored in a storage container accessible via a server device, the MPI streaming content including a plurality of coded bitstreams that includes the first coded bitstream; for a period, providing to the client device a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; Docket No. D23136WO01 receiving, from the client device, a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and transmitting to the client device one or more of the plurality of coded bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. [0134] In some embodiments of any of the above methods, for the period, the storage container has a plurality of media segments logically organized in accordance with different views and further logically organized in accordance with one or more of different bit rates, different resolutions, different codec types, and different frame rates. [0135] In some embodiments of any of the above methods, in a multi-track mode, a common atlas, an individual atlas, and a packed video component are represented in the MPD file as separate adaptation sets; and wherein the adaptation set for the common atlas contains the respective initialization segment having parameter sets for initializing a V3C decoder at the client device. [0136] In some embodiments of any of the above methods, media segments for a representation of the adaptation set for the common atlas contain one or more track fragments of the V3C atlas track; and wherein media segments for representations of the video component adaptation sets contain one or more track fragments of the corresponding packed video track. [0137] In some embodiments of any of the above methods, the first coded bitstream is encapsulated in a single V3C bitstream track stored in the storage container. [0138] In some embodiments of any of the above methods, the first coded bitstream is encapsulated in two or more V3C bitstream tracks stored in the storage container. [0139] In some embodiments of any of the above methods, the method further comprises switching from a first sequence of media segments to a different second sequence of media segments when the request indicates a change in the identified selection. [0140] In some embodiments of any of the above methods, the transmitting includes: transmitting a first coded bitstream carrying the first sequence of media segments corresponding to a first one of the different views; and transmitting a second coded bitstream carrying the different second sequence of media segments corresponding to a second one of the different views, wherein the first coded bitstream and the second coded bitstream have respective media segments corresponding to a same one of the different respective video segment times. [0141] In some embodiments of any of the above methods, the MPI video is a multiview MPI video. Docket No. D23136WO01 [0142] In some embodiments of any of the above methods, the method further comprises: providing to a client device a media presentation description corresponding to the multiview MPI video; for a period, providing to the client device a respective initialization segment to inform a selection, at the client device, of two or more camera views of the multiview MPI video for which to request media segments for rendering; receiving, from the client device, a request identifying the selection; and transmitting to the client device one or more coded bitstreams carrying media segments in accordance with the identified selection. [0143] In some embodiments of any of the above methods, the method further comprises switching from transmitting a first sequence of media segments to transmitting a different second sequence of media segments when the request indicates a change of at least one camera view in the identified selection of the two or more camera views. [0144] In some embodiments of any of the above methods, the transmitting includes: transmitting a first sequence of media segments corresponding to a first camera view of a scene; and transmitting a different second sequence of media segments corresponding to a different second camera view of the scene, wherein the first and second sequences of media segments correspond to a same time interval of the multiview MPI video. [0145] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods. [0146] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims. [0147] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated Docket No. D23136WO01 into such future embodiments. In sum, it should be understood that the application is capable of modification and variation. [0148] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary. [0149] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter. [0150] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims. [0151] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit. [0152] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and/or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for Docket No. D23136WO01 practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits. All pseudo code examples presented herein are strictly for illustration purposes that illustrate applicable coding concepts that can be implemented in various components of the disclosed MPI video system. Based on the provided examples, a person of ordinary skill in the pertinent art will be able to make and use functionally similar blocks of code for various specific implementations of the corresponding MPI video system without any undue experimentation. [0153] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range. [0154] The use of figure numbers and/or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures. [0155] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence. [0156] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.” [0157] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner. [0158] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated Docket No. D23136WO01 condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].” [0159] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements. [0160] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard. [0161] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and/or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and/or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context. [0162] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and/or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and/or digital hardware circuit(s) with software/firmware and (ii) any portions of hardware processor(s) with software Docket No. D23136WO01 (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. [0163] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown. [0164] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and/or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

Claims

Docket No. D23136WO01 CLAIMS What is claimed is: 1. A method for streaming a multiplane-image (MPI) video, the method comprising: generating a sequence of video frames, each of the video frames including a respective plurality of patches representing texture and transparency layers of one or more multiplane images of the MPI video; applying video compression to the sequence of video frames to generate a video sub- stream; generating a sequence of representations of atlas frames corresponding to the sequence of video frames to specify at least a packing arrangement of the patches; applying compression to the sequence of representations to generate an atlas sub-stream; and multiplexing the video sub-stream and the atlas sub-stream to generate a first coded bitstream encoding at least a portion of the MPI video. 2. The method of claim 1, wherein the first coded bitstream is a visual volumetric video- based coding (V3C) bitstream in accordance with the ISO.IEC.23090-5 standard. 3. The method of claim 2, further comprising encapsulating the first coded bitstream into a media file configured for file playback. 4. The method of claim 2, further comprising: encapsulating the first coded bitstream into a sequence of segments according to a selected media container file format; and streaming the sequence of segments over a communication channel to a client device for playback. 5. The method of claim 2, wherein the first coded bitstream is configured for encapsulation in a single V3C bitstream track. 6. The method of claim 2, wherein the first coded bitstream is configured for encapsulation in two or more V3C bitstream tracks. Docket No. D23136WO01 7. The method of claim 6, wherein the two or more V3C bitstream tracks include: a first track that carries at least a portion of the atlas sub-stream; and a second track that carries at least a portion of the video sub-stream. 8. The method of claim 1, wherein texture and transparency patches of the respective plurality of patches are packed into a corresponding video frame using a packing arrangement selected from the group consisting of: a side-by-side arrangement; a top-to-bottom arrangement; a vertically interleaved arrangement; and a horizontally interleaved arrangement. 9. The method of claim 8, wherein the texture and transparency patches include patches of a first size and patches of a different second size. 10. The method of claim 1, wherein the respective plurality of patches has fewer than all of the patches of a corresponding multiplane image. 11. The method of claim 1, further comprising: providing to a client device a media presentation description (MPD) of an MPI streaming content stored in a storage container accessible via a server device, the MPI streaming content including a plurality of coded bitstreams that includes the first coded bitstream; for a period, providing to the client device a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; receiving, from the client device, a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and transmitting to the client device one or more of the plurality of coded bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. Docket No. D23136WO01 12. The method of claim 11, wherein, for the period, the storage container has a plurality of media segments logically organized in accordance with different views and further logically organized in accordance with one or more of different bit rates, different resolutions, different codec types, and different frame rates. 13. The method of claim 11, wherein, in a multi-track mode, a common atlas, an individual atlas, and a packed video component are represented in the MPD file as separate adaptation sets; and wherein the adaptation set for the common atlas contains the respective initialization segment having parameter sets for initializing a V3C decoder at the client device. 14. The method of claim 13, wherein media segments for a representation of the adaptation set for the common atlas contain one or more track fragments of the V3C atlas track; and wherein media segments for representations of the video component adaptation sets contain one or more track fragments of the corresponding packed video track. 15. The method of claim 11, wherein the first coded bitstream is encapsulated in a single V3C bitstream track stored in the storage container. 16. The method of claim 11, wherein the first coded bitstream is encapsulated in two or more V3C bitstream tracks stored in the storage container. 17. The method of claim 11, further comprising switching from a first sequence of media segments to a different second sequence of media segments when the request indicates a change in the identified selection. 18. The method of claim 17, wherein the transmitting includes: transmitting a first coded bitstream carrying the first sequence of media segments corresponding to a first one of the different views; and transmitting a second coded bitstream carrying the different second sequence of media segments corresponding to a second one of the different views, wherein the first coded bitstream and the second coded bitstream have respective media segments corresponding to a same one of the different respective video segment times. Docket No. D23136WO01 19. The method of claim 1, wherein the MPI video is a multiview MPI video. 20. The method of claim 19, further comprising: providing to a client device a media presentation description corresponding to the multiview MPI video; for a period, providing to the client device a respective initialization segment to inform a selection, at the client device, of two or more camera views of the multiview MPI video for which to request media segments for rendering; receiving, from the client device, a request identifying the selection; and transmitting to the client device one or more coded bitstreams carrying media segments in accordance with the identified selection. 21. The method of claim 20, further comprising switching from transmitting a first sequence of media segments to transmitting a different second sequence of media segments when the request indicates a change of at least one camera view in the identified selection of the two or more camera views. 22. The method of claim 20, wherein the transmitting includes: transmitting a first sequence of media segments corresponding to a first camera view of a scene; and transmitting a different second sequence of media segments corresponding to a different second camera view of the scene, wherein the first and second sequences of media segments correspond to a same time interval of the multiview MPI video. 23. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of claim 1. 24. An apparatus for streaming a multiplane-image (MPI) video, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: Docket No. D23136WO01 generate a sequence of video frames, each of the video frames including a respective plurality of patches representing texture and transparency layers of one or more multiplane images of the MPI video; apply video compression to the sequence of video frames to generate a video sub-stream; generate a sequence of representations of atlas frames corresponding to the sequence of video frames to specify at least a packing arrangement of the patches; apply compression to the sequence of representations to generate an atlas sub-stream; and multiplex the video sub-stream and the atlas sub-stream to generate a first coded bitstream encoding at least a portion of the MPI video.
EP24743159.6A 2023-06-27 2024-06-24 Multi-view multiplane-imaging video streaming Pending EP4736458A1 (en)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
US202363510571P 2023-06-27 2023-06-27
US202363588337P 2023-10-06 2023-10-06
US202463556461P 2024-02-22 2024-02-22
PCT/US2024/035239 WO2025006382A1 (en) 2023-06-27 2024-06-24 Multi-view multiplane-imaging video streaming

Publications (1)

Publication Number Publication Date
EP4736458A1 true EP4736458A1 (en) 2026-05-06

Family

ID=91950290

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24743159.6A Pending EP4736458A1 (en) 2023-06-27 2024-06-24 Multi-view multiplane-imaging video streaming

Country Status (4)

Country Link
EP (1) EP4736458A1 (en)
CN (1) CN121569487A (en)
MX (2) MX2025015607A (en)
WO (1) WO2025006382A1 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP7649792B2 (en) * 2020-04-15 2025-03-21 中興通訊股▲ふん▼有限公司 Volumetric visual media processing method and apparatus
CN117413521A (en) * 2021-05-06 2024-01-16 交互数字Ce专利控股有限公司 Method and apparatus for encoding/decoding a volumetric video, and method and apparatus for reconstructing a computer generated hologram

Also Published As

Publication number Publication date
WO2025006382A1 (en) 2025-01-02
CN121569487A (en) 2026-02-24
MX2025015607A (en) 2026-03-02
MX2026000850A (en) 2026-03-02

Similar Documents

Publication Publication Date Title
CN115211131B (en) Apparatus, method and computer program for omni-directional video
CN109076255B (en) Method and device for sending and receiving 360-degree video
CN111164969B (en) Method and apparatus for sending or receiving 6DOF video using stitching and reprojection related metadata
KR102559862B1 (en) Methods, devices, and computer programs for media content transmission
CN108702528B (en) Method for sending 360 video, method for receiving 360 video, device for sending 360 video and device for receiving 360 video
CN111819842B (en) Method and device for sending 360-degree video, and method and device for receiving 360-degree video
CN110870321B (en) Region-wise packaging, content coverage, and signaling frame packaging for media content
KR102221301B1 (en) Method and apparatus for transmitting and receiving 360-degree video including camera lens information
US11051040B2 (en) Method and apparatus for presenting VR media beyond omnidirectional media
CN110800311B (en) Method, apparatus and computer program for transmitting media content
CN111727605B (en) Method and apparatus for transmitting and receiving metadata regarding multiple viewpoints
KR102214079B1 (en) Method for transmitting 360-degree video, method for receiving 360-degree video, apparatus for transmitting 360-degree video, and apparatus for receiving 360-degree video
AU2018271975A1 (en) High-level signalling for fisheye video data
EP3873095A1 (en) An apparatus, a method and a computer program for omnidirectional video
EP4736458A1 (en) Multi-view multiplane-imaging video streaming
EP4736460A1 (en) Architecture for interactive multi-view multiplane-imaging video streaming

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE