EP4736460A1 - Architecture for interactive multi-view multiplane-imaging video streaming - Google Patents

Architecture for interactive multi-view multiplane-imaging video streaming

Info

Publication number
EP4736460A1
EP4736460A1 EP24739985.0A EP24739985A EP4736460A1 EP 4736460 A1 EP4736460 A1 EP 4736460A1 EP 24739985 A EP24739985 A EP 24739985A EP 4736460 A1 EP4736460 A1 EP 4736460A1
Authority
EP
European Patent Office
Prior art keywords
mpi
different
information
storage container
media segments
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24739985.0A
Other languages
German (de)
French (fr)
Inventor
Peng Yin
Guan-Ming Su
Taoran Lu
Dae Yeol Lee
Tsung-Wei Huang
Sean Thomas MCCARTHY
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Dolby Laboratories Licensing Corp
Original Assignee
Dolby Laboratories Licensing Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Dolby Laboratories Licensing Corp filed Critical Dolby Laboratories Licensing Corp
Publication of EP4736460A1 publication Critical patent/EP4736460A1/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/21Server components or server architectures
    • H04N21/218Source of audio or video content, e.g. local disk arrays
    • H04N21/21805Source of audio or video content, e.g. local disk arrays enabling multiple viewpoints, e.g. using a plurality of cameras
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/234Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
    • H04N21/2343Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements
    • H04N21/23439Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements for generating different versions
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/20Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
    • H04N21/23Processing of content or additional data; Elementary server operations; Server middleware
    • H04N21/235Processing of additional data, e.g. scrambling of additional data or processing content descriptors
    • H04N21/2353Processing of additional data, e.g. scrambling of additional data or processing content descriptors specifically adapted to content descriptors, e.g. coding, compressing or processing of metadata
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/60Network structure or processes for video distribution between server and client or between remote clients; Control signalling between clients, server and network components; Transmission of management data between server and client, e.g. sending from server to client commands for recording incoming content stream; Communication details between server and client 
    • H04N21/65Transmission of management data between client and server
    • H04N21/658Transmission by the client directed to the server
    • H04N21/6587Control parameters, e.g. trick play commands, viewpoint selection
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/81Monomedia components thereof
    • H04N21/816Monomedia components thereof involving special video data, e.g 3D video
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/84Generation or processing of descriptive data, e.g. content descriptors
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/845Structuring of content, e.g. decomposing content into time segments
    • H04N21/8456Structuring of content, e.g. decomposing content into time segments by decomposing the content in the time domain, e.g. in time segments

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Library & Information Science (AREA)
  • Databases & Information Systems (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)

Abstract

Methods and apparatus for multi-plane-image (MPI) video streaming. According to an example embodiment, a method of video streaming includes providing to a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device and a respective initialization segment from the storage container. The respective initialization segment is configured to inform a selection of views of the MPI streaming content for which to request media segments for rendering. The method also includes receiving, from the client device, a request identifying the selection and indicating a respective recommended value of at least one of a bit rate, a resolution, a codec type, and a frame rate and transmitting to the client device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values.

Description

Docket No.: D23069WO01 ARCHITECTURE FOR INTERACTIVE MULTI-VIEW MULTIPLANE-IMAGING VIDEO STREAMING 1. Cross-Reference to Related Applications [0001] This application claims the priority benefit of U.S. Provisional Patent Application No. 63/510,571 filed June 27, 2023, the contents of which are incorporated by reference in its entirety. 2. Field of the Disclosure [0002] Various example embodiments relate generally to multiplane imaging (MPI) and, more specifically but not exclusively, to transmission of multiplane images. 3. Background [0003] Multiplane images embody a relatively new approach to storing volumetric content. MPI can be used to render both still images and video and represents a three-dimensional (3D) scene within a view frustum using, e.g., 8, 16, or 32 planes of texture and transparency (alpha) information per camera. Example applications of MPI include computer vision and graphics, image editing, photo animation, robotics, and virtual reality. BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS [0004] Disclosed herein are various embodiments of an end-to-end interactive MPI video- streaming system. Various examples of MPI storage format and transport protocol used in the MPI video-streaming system are described, including descriptions of the Media Presentation Description (MPD), initialization segment, and media segment. At least some embodiments of the MPI video-streaming system provide seamless interactive view switching for multi-view MPI video-streaming experience. Some embodiments are directed at providing interactive immersive volumetric experience by leveraging pertinent components of available technologies, such as the 2D video codec, MPEG Dynamic Adaptive Streaming over HTTP (DASH), and ISO Base Media File Format (BMFF) and its extensions. At least some of the disclosed solutions can be deployed in a relatively short time, after modifications of pertinent existing solutions are implemented in accordance with various embodiments disclosed herein. [0005] According to an example embodiment, provided is a method for MPI video streaming comprising: providing to a client device a media presentation description of an MPI streaming Docket No.: D23069WO01 content stored in a storage container accessible via a server device; for a period, providing to the client device a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; receiving, from the client device, a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and transmitting to the client device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. [0006] According to another example embodiment, provided is a non-transitory computer- readable medium storing instructions that, when executed by an electronic processor of a server device, cause the server device to perform operations comprising the above method for MPI video streaming. [0007] According to yet another example embodiment, provided is an apparatus for MPI video streaming, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: provide to a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, provide to the client device a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; receive, from the client device, a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and transmit to the client device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended parameter values. [0008] According to yet another example embodiment, provided is a method for MPI video streaming, comprising: receiving at a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, receiving at the client device, via the server device, a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media Docket No.: D23069WO01 segments for rendering; transmitting to the server device a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and receiving from the server device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. [0009] According to yet another example embodiment, provided is a non-transitory computer-readable medium storing instructions that, when executed by an electronic processor of a client device, cause the client device to perform operations comprising the above method for MPI video streaming. [0010] According to yet another example embodiment, provided is an apparatus for MPI video streaming, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: receive at a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, receive at the client device, via the server device, a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; transmit to the server device a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and receive from the server device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. BRIEF DESCRIPTION OF THE DRAWINGS [0011] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which: [0012] FIG.1 depicts an example process for a video/image delivery pipeline. [0013] FIG.2 pictorially illustrates a 3D-scene representation using a multiplane image according to an embodiment. Docket No.: D23069WO01 [0014] FIG.3 pictorially illustrates a process of generating a novel view of a 3D scene according to one example. [0015] FIG.4 is a block diagram illustrating a change of the set of active views over time according to one example. [0016] FIG.5 is a block diagram of an example communication system in which MPEG DASH streaming can be implemented according to some embodiments. [0017] FIG.6 is a block diagram illustrating a high-level hierarchical data model that can be used in DASH according to some embodiments. [0018] FIG.7 is a block diagram illustrating example practical applications assigned to various hierarchical levels of the data model of FIG.6 according to some embodiments. [0019] FIG.8 is a pseudocode illustrating a common tag using which a Media Presentation Description (MPD) is described in the XML format according to some examples. [0020] FIG.9 shows an example of a restricted video sample entry that can be used in some embodiments. [0021] FIG.10 is a block diagram illustrating an overall structure of an ISO BMFF file for timed data according to some examples. [0022] FIG.11 is a block diagram illustrating an overall segment structure of fragmented MP4 according to some examples. [0023] FIG.12 is a block diagram illustrating an overall structure of a DASH representation according to some examples. [0024] FIG.13 is a block diagram illustrating a communication system configured to support MPI transmissions using DASH and ISO BMFF according to some embodiments. [0025] FIG.14 is a block diagram illustrating a data structure of an MPI container used in the communication system of FIG.13 according to some embodiments. [0026] FIG.15 shows an example of the MPD that can be used in the communication system of FIG.13 according to some embodiments. Docket No.: D23069WO01 [0027] FIG.16 shows another example of the MPD that can be used in the communication system of FIG.13 according to some embodiments. [0028] FIG.17 shows yet another example of the MPD that can be used in the communication system of FIG.13 according to some embodiments. [0029] FIG.18 shows still another example of the MPD that can be used in the communication system of FIG.13 according to some embodiments. [0030] FIGS.19A-19F provide example definitions related to a view identifier used in the communication system of FIG.13 according to some embodiments. [0031] FIGS.20A and 20B provide example definitions related to the intrinsic camera parameters used in the communication system of FIG.13 according to some embodiments. [0032] FIGS.21A and 21B provide example definitions related to the extrinsic camera parameters used in the communication system of FIG.13 according to some embodiments. [0033] FIG.22 provides example definitions related to a restricted scheme information used in the communication system of FIG.13 according to some embodiments. [0034] FIGS.23A and 23B provide example definitions related to the SEI information box used in the communication system of FIG.13 according to some embodiments. [0035] FIG.24 shows an example of a method for MPI video arrangement used in the communication system of FIG.13 according to some embodiments. [0036] FIG.25 provides example definitions related to MPI video arrangements used in the communication system of FIG.13 according to some embodiments. [0037] FIG.26 shows an example of another method for MPI video arrangement used in the communication system of FIG.13 according to some embodiments. [0038] FIG.27 shows an example used to specify a multi-view camera setup in the communication system of FIG.13 according to some embodiments. [0039] FIGS.28A-28C provide example definitions related pre-fetch operations used in the communication system of FIG.13 according to some embodiments. Docket No.: D23069WO01 [0040] FIG.29 is a block diagram illustrating encapsulation of multi-view MPI coded data in one ISO BMFF file used for storage according to some embodiments. [0041] FIG.30 is a block diagram illustrating a computing device used in the communication system of FIG.13 according to an embodiment. DETAILED DESCRIPTION [0042] This disclosure and aspects thereof can be embodied in various forms, including hardware, devices or circuits controlled by computer-implemented methods, computer program products, computer systems and networks, user interfaces, and application programming interfaces; as well as hardware-implemented methods, signal processing circuits, memory arrays, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), and the like. The foregoing is intended solely to give a general idea of various aspects of the present disclosure and does not limit the scope of the disclosure in any way. [0043] In the following description, numerous details are set forth, such as device configurations, timings, operations, and the like, in order to provide an understanding of one or more aspects of the present disclosure. It will be readily apparent to one skilled in the art that these specific details are merely exemplary and not intended to limit the scope of this application. Example Video/Image Delivery Pipeline [0044] FIG.1 depicts an example process of a video delivery pipeline (100), showing various stages from video/image capture to video/image-content display according to an embodiment. A sequence of video/image frames (102) may be captured or generated using an image-generation block (105). The frames (102) may be digitally captured (e.g., by a digital camera) or generated by a computer (e.g., using computer animation) to provide video and/or image data (107). Alternatively, the frames (102) may be captured on film by a film camera. Then, the film may be translated into a digital format to provide the video/image data (107). In some examples, the image-generation block (105) includes generating an MPI image or video. [0045] In a production phase (110), the data (107) may be edited to provide a video/image production stream (112). The data of the video/image production stream (112) may be provided to a processor (or one or more processors, such as a central processing unit, CPU) at a post- production block (115) for post-production editing. The post-production editing of the block Docket No.: D23069WO01 (115) may include, e.g., adjusting or modifying colors or brightness in particular areas of an image to enhance the image quality or achieve a particular appearance for the image in accordance with the video creator’s creative intent. This part of post-production editing is sometimes referred to as “color timing” or “color grading.” Other editing (e.g., scene selection and sequencing, image cropping, addition of computer-generated visual special effects, removal of artifacts, etc.) may be performed at the block (115) to yield a “final” version (117) of the production for distribution. In some examples, operations performed at the block (115) include enhancing texture and/or alpha channels in multiplane images/video. During the post-production editing (115), video and/or images may be viewed on a reference display (125). [0046] Following the post-production (115), the data of the final version (117) may be delivered to a coding block (120) for being further delivered downstream to decoding and playback devices, such as television sets, set-top boxes, movie theaters, and the like. In some embodiments, the coding block (120) may include audio and video encoders, such as those defined by the ATSC, DVB, DVD, Blu-Ray, and other delivery formats, to generate a coded bitstream (122). In a receiver, the coded bitstream (122) is decoded by a decoding unit (130) to generate a corresponding decoded signal (132) representing a copy or a close approximation of the signal (117). The receiver may be attached to a target display (140) that may have somewhat or completely different characteristics than the reference display (125). In such cases, a display management (DM) block (135) may be used to map the decoded signal (132) to the characteristics of the target display (140) by generating a display-mapped signal (137). Depending on the embodiment, the decoding unit (130) and display management block (135) may include individual processors or may be based on a single integrated processing unit. [0047] A codec used in the coding block (120) and/or the decoding block (130) enables video/image data processing and compression/decompression. The compression is used in the coding block (120) to make the corresponding file(s) or stream(s) smaller. The decoding process carried out by the decoding block (130) typically includes decompressing the received video/image data file(s) or streams(s) into a form usable for playback and/or further editing. Example coding/decoding operations that can be used in the coding block (120) and the decoding unit (130) according to various embodiments are described in more details below. Multiplane Imaging [0048] A multiplane image comprises multiple image planes, with each of the image planes being a “snapshot” of the 3D scene at a certain depth with respect to the camera position. Docket No.: D23069WO01 Information stored in each plane includes the texture information (e.g., represented by the R, G, B values) and transparency information (e.g., represented by the alpha (A) values). Herein, the acronyms R, G, B stand for red, green, and blue, respectively. In some examples, the three texture components can be (Y, Cb, Cr), or (I, Ct, Cp), or another functionally similar set of values. There are different ways in which a multiplane image can be generated. For example, two or more input images from two or more cameras located at different known viewpoints can be co-processed to generate a corresponding multiplane image. Alternatively, a multiplane image can be generated using a source image captured by a single camera. [0049] FIG.2 pictorially illustrates a 3D scene representation using a multiplane image (200) according to an embodiment. The multiplane image (200) has D planes or layers (P0, P1, …, P(D−1)), where D is an integer greater than one. Typically, the planes (layers) are indexed such that the most remote layer, from the reference camera position (RCP), is indexed as the 0-th layer and is at a distance (or depth) d0 from the RCP along the Z dimension of the 3D scene. The index is incremented by one for each next layer located closer to the RCP. The plane (layer) that is the closest to the RCP has the index value (D−1) and is at a distance (or depth) dD−1 from the RCP along the Z dimension. Each of the planes (P0, P1, …, P(D−1)) is orthogonal to a base plane (202) which is parallel to the XZ-coordinate plane. The RCP is at a vertical height h above the base plane (202). The XYZ triad shown in FIG.2 indicates the general orientation of the multiplane image (200) and the planes (P0, P1, …, P(D−1)) with respect to the X, Y, and Z dimensions of the 3D scene. In various examples, the number D can be 32, 16, 8, or any other suitable integer greater than one. [0050] Let us denote the color component (e.g., RGB) value for the ith layer at camera location s as ^^^^( ^^^^) ^^^^ , with the lateral size of the layer being H×W, where H is the height (Y dimension) and W is the width (X dimension) of the layer. The pixel value at location (x, y) for the color channel c is represented as ^^^^( ^^^^) th ( ^^^^) ^^^^ ( ^^^^, ^^^^, ^^^^). The α value for the i layer is ^^^^ ^^^^ . The pixel value (x, y) in the alpha layer is represented as ^^^^( ^^^^)( th ^^^^ ^^^^, ^^^^). The depth distance between the i layer to the reference camera position is ^^^^ ^^^^. The image from the original reference view (without the camera moving) is denoted as ^^^^, with the texture pixel value being ^^^^( ^^^^)( ^^^^, ^^^^, ^^^^). A still MPI image for the camera location s can therefore be represented as: Docket No.: D23069WO01 It is straightforward to extend this still MPI image representation to a video representation, provided that the camera position s is kept static overtime. This video representation is given by Eq. (2): where t denotes time. [0051] As already indicated above, a multiplane image, such as the multiplane image (200), can be generated from a single source image ^^^^ or from two or more source images. Such generation may be performed, e.g., during the production phase (110). The corresponding MPI generation algorithm(s) may typically output the multiplane image (200) containing XYZ- resolved pixel values in the form {( ^^^^ ^^^^, ^^^^ ^^^^) for i=0, …, D−1}. [0052] By processing the multiplane image (200) represented by {( ^^^^ ^^^^, ^^^^ ^^^^) for i=0, …, D−1}, an MPI-rendering algorithm can generate a viewable image corresponding to the RCP or to a new virtual camera position that is different from the RCP. An example MPI-rendering algorithm (often referred to as the “MPI viewer”) that can be used for this purpose may include the steps of warping and compositing. Other suitable MPI viewers may also be used. The rendered multiplane image (200) can be viewed, e.g., on the reference display (125). [0053] During the warping step of the MPI-rendering algorithm, each layer ( ^^^^ ^^^^, ^^^^ ^^^^) of the multiplane image (200) may be warped from the RCP viewpoint position ( ^^^^ ^^^^) to a new viewpoint position ( ^^^^ ^^^^), e.g., as follows: ^^^^ ^ ^^ ^ ^^ ^^ = ^^^^ ^^^^ ^^^^, ^^^^ ^^^^( ^^^^ ^^^^ ^^^^ , ^^^^ ^^^^) (3) where ^^^^ ^^^^ ^^^^, ^^^^ ^^^^() is the warping function; and σ is the consistent scale (to minimize error). In an example embodiment, the warping function ^^^^ ^^^^ ^^^^, ^^^^ ^^^^() can be expressed as follows: where ^^^^ ^^^^ = ( ^^^^ ^^^^, ^^^^ ^^^^) and ^^^^ ^^^^ = ( ^^^^ ^^^^ , ^^^^ ^^^^). Through (5), each pixel location ( ^^^^ ^^^^ , ^^^^ ^^^^) on the target view of a certain MPI plane can be mapped to its respective pixel location ( ^^^^ ^^^^, ^^^^ ^^^^) on the source view. The functions ^^^^ ^^^^ and ^^^^ ^^^^ represent the intrinsic camera model for the reference view and the target view, respectively. The functions R and t represent the extrinsic camera model for rotation and translation, respectively. n denotes the normal vector [001]T. a denotes the distance to a plane that is fronto-parallel to the source camera at depth ^^^^ ^^^^ ^^^^. Docket No.: D23069WO01 [0054] During the compositing step of the MPI-rendering algorithm, a new viewable image ^^^^ ^^^^ can be generated, e.g., using processing operations corresponding to the following equations: = ∑ ^ ^^ ^ ^^ ^ =^− 01 ^^^^ ^ ^^ ^ ^^ ^^ ^^^^ ^ ^^ ^ ^^ ^^ (6) where the weights ^^^^ ^ ^^ ^ ^^ ^^ are expressed as: ^^^^ ^ ^^ ^ ^^ ^^ = ( ^^^^ ^ ^^ ^ ^^ ^^ ∙ ∏ ^ ^^ ^ ^^ ^^ = ^^^ 1 ^+1 (1 − ^^^^ ^ ^ ^^ ^^ ^^) ) (7) The disparity map ^^^^ ^^^^ corresponding to the source view can be computed as: ^^^^ ^^^^ = ∑ ^ ^^ ^ ^^ ^ =^− 01 ^^^^ −1 ^^^^ ^^^^ ^ ^ ^^ ^^ ^^ (8) where the weights ^^^^ ^ ^ ^^ ^^ ^^ are expressed as: The MPI-rendering algorithm can also be used to generate the viewable image ^^^^ ^^^^ corresponding to the RCP. In this case, the warping step is omitted, and the image ^^^^ ^^^^ is computed as: ^^^^ ^^^^ = ∑ ^ ^^ ^ ^^ ^ =^− 01 ^^^^ ^ ^ ^^ ^^ ^^ ^^^^ ^ ^ ^^ ^^ ^^ (10) [0055] In the single camera transmission scenario, only one MPI is fed through a bitstream. A goal for this situation is to optimally merge the layers of the original MPI such that the quality of this MPI after local warping is preserved. In the multiple camera transmission scenario, multiple MPIs captured in different camera positions are encoded in the compressed bitstream. The information in these MPIs is jointly used to generate global novel views for positions located between the original camera positions. There also can be a scenario where information from multiple cameras can be used jointly to generate a single MPI to be transmitted. For transmissions of MPI video, the multiple camera transmission scenario is typically used, e.g., as explained below. [0056] FIG.3 pictorially illustrates a process of generating a novel view of a 3D scene (302) according to one example. In the example shown, the 3D scene (302) is captured using forty- two RCPs (1, 2, …, 42). The novel view that is being generated corresponds to a camera position (50). The four closest RCPs to the camera position (50) are the RCPs (11, 12, 18, 19). The corresponding multiplane images are multiplane images (20011, 20012, 20018, 20019). A multiplane image (20050) corresponding to the camera position (50) is generated by correspondingly warping the multiplane images (20011, 20012, 20018, 20019) and then merging the resulting warped multiplane images. Finally, a viewable image (312) of the 3D scene (302) is generated by applying the compositing step (310) of the MPI-rendering algorithm to the multiplane image (20050). Docket No.: D23069WO01 [0057] In general, a 3D scene, such as the 3D scene (302) may be captured using any suitably selected number of RCPs. The locations of such RCPs can also be variously selected, e.g., based on the creative intent. In typical practical examples, when a novel view, such as the viewable image (312) is rendered, only several neighboring RCPs are used for the rendering. Hereafter, such neighboring views are referred to as the “active views.” In the example illustrated in FIG.3, the number of active views is four. In other examples, a different (from four) number of active views may similarly be used. As such, the number of active views is a selectable parameter. For illustration purposes and without any implied limitations, some example embodiments are described herein below in reference to four active views. In some examples, the set of active views may change over time when the camera position (50) moves. In some examples, the number of active views may change over time when the camera position (50) moves. [0058] FIG.4 is a block diagram illustrating a change of the set of active views over time according to one example. In the example shown, a 3D scene is captured using a rectangular array of forty RCPs arranged in five rows and eight columns. A dashed arrow (402) represents a movement trajectory of the novel camera position (50) during the time interval starting at the time t0 and ending at the time t1 for virtual view synthesis. At the time t0, the set of active views includes the four views encompassed by the dashed box (410). At the time t1, the set of active views includes the four views encompassed by the dotted box (420). At the time tn (where t0< tn < t1), the set of active views changes from being the set in the dashed box (410) to being the set in the dotted box (420). [0059] Various embodiments disclosed herein are directed at providing an end-to-end interactive multi-view MPI video streaming system including: • Components located at the server side and content storage formats used therein; • Components supporting through-the-network distribution protocols; and • Components supporting the end client-side interactive playback solutions. Some embodiments use the MPEG-DASH and ISO-BMFF standards and/or their extensions and additionally provide new tools that support new functionalities. One goal of such embodiments is to enable seamless interactive view switching for multi-view MPI video streaming by building up on and expanding certain legacy technologies/infrastructures to enable expedited and/or cost- efficient deployment of the interactive immersive volumetric video experience. Some embodiments may benefit from at least some features disclosed in one or more of the following documents: (1) ISO/IEC 14496-12:2020, Information technology -- Coding of audio-visual Docket No.: D23069WO01 objects -- Part 12: ISO base media file format; (2) ISO/IEC 14496-15:2021, Information technology -- Coding of audio-visual objects -- Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format; (3) ISO/IEC 23000-19:2018, Information technology — Multimedia application format (MPEG-A) -- Part 19: Common media application format (CMAF) for segmented media, and (4) ISO/IEC 23009-1(2019), Information Technology -- Dynamic Adaptive Streaming over HTTP (DASH) -- Part 1: Media Presentation Description and Segment Formats, . Each of these documents is incorporated herein by reference in its entirety. [0060] In at least some examples, multi-view MPI representations may convey one or both of the following benefits: • Improve single MPI novel-view spatial rendering. In some examples, up to four views are sufficient. In some additional examples, with a different MPI generation process, the number of views can be reduced, e.g., down to a single view. • Enable novel view rendering both spatially and temporally over time, e.g., to provide a greater choice of render poses and a better experience. An example of spatial and temporal switching of reference views is described above in reference to FIG.4. [0061] In some examples, a multi-view MPI system operates to independently code every single-view MPI using one of the frame-packing solutions disclosed in commonly owned U.S. Patent Application Serial No.63/495,715, filed on 12-APR-2023 and entitled “TRANSMISSION OF VOLUMETRIC IMAGES IN MULTIPLANE IMAGING FORMAT,” which is incorporated herein by reference in its entirety. In some use cases, a user will request a different novel view pose at different times. Accordingly, the client side of the system needs to have a capability to dynamically switch to different MPI bitstreams. One solution is to have all views’ MPI bitstreams downloaded to the client side. However, this solution may not be bandwidth efficient. Another solution is to stream a smaller number of appropriately selected views’ MPI bitstreams to improve the network and/or computing resource utilization. Example challenges to practically implementing the latter solution include: • How to efficiently store MPI video bitstreams in the serve side; • How to efficiently transmit MPI video bitstreams from the server side to the client side; and • How to efficiently manage the MPI bitstream switching and decode/render the novel view according to the interactivity from the users and/or network. Docket No.: D23069WO01 [0062] Some embodiments disclosed herein below provide solutions for building an end-to- end interactive multi-view MPI video streaming system. For the storage format, various embodiments build on the ISO base media file format, commonly referred to as ISO BMFF, (ISO/IEC 14496-12), and its extension ISO/IEC 14496-15, which provides carriage of the network abstraction layer (NAL) unit structured video, such as the AVC, HEVC, VVC, and the like. Herein, the acronym ISO stands for the International Organization for Standardization. For the data transport, various embodiments built on Dynamic adaptive streaming over HTTP (DASH) (ISO/IEC 230090-1). In some examples, a similar approach is applied to the Common Media Application Format (CMAF). Note that the disclosed video streaming system can also support other multi-view coding solutions, such as the MVC, SHVC, and multi-layer VVC. For illustration purposes and without any implied limitations, example embodiments are described in reference to the HEVC codec. Based on the provided description, a person of ordinary skill in the pertinent art will be able to make and use other embodiments compatible with other suitable codecs without any undue experimentation. MPEG DASH [0063] FIG.5 is a block diagram of an example communication system (500) in which MPEG DASH streaming can be implemented according to some embodiments. MPEG DASH is a media streaming technology that enables high-quality adaptive streaming of media content (such as, video, audio, etc.) over the Internet delivered through HTTP web servers and proxies. In the example shown, the communication system (500) includes a media capture device (510) connected to a computation device (520). The computation device (520) is further network connected to a plurality of servers including a Digital Rights Management (DRM) encryption server (530), media origin servers (540), and HTTP cache servers (550). A plurality of client devices (560) is connectable to receive digital content from the HTTP cache servers (550), provided that a proper license for the digital content can be obtained from one of DRM license servers (570). [0064] An example of DASH-streaming procedure from the client’s side prospective includes the following blocks of operations: 1) A client device (560) obtains the Media Presentation Description (MPD) of a streaming content, e.g., a movie. In some examples, the MPD includes information on different alternative representations, e.g., bit rate, video resolution, frame rate, audio language of the streaming content, as well as the URLs of the HTTP resources (the initialization segment and the media segments). Docket No.: D23069WO01 2) The client device (560) requests one or more of the desired representations, one segment (or a part thereof) at a time, based on pertinent information in the MPD and the client’s local information, e.g., network bandwidth, decoding/display capabilities, and/or user preferences. 3) The client device (560) requests segments of a different representation with a better- matching bitrate, e.g., when the client device detects a network bandwidth change. In some cases, it starts from a segment that starts with a random-access point. [0065] FIG.6 is a block diagram illustrating a high-level hierarchical data model (600) that can be used in DASH according to some embodiments. In one dimension (illustratively the horizontal domain), the model (600) includes a sequence in time of the corresponding Media Presentation. In another dimension (illustratively the vertical domain), the model (600) includes the choices offered in the Media Presentation, to be selected by the DASH Client in a static or dynamic manner. [0066] In some examples, the model (600) relies on the following definition of various levels used in the DASH hierarchy: • Media Presentation Description (MPD): a DASH Media Presentation is described by an MPD document (602). This document describes a sequence of Periods (610) in time that make up the Media Presentation. • Period: a period (610) typically represents a media content period during which a consistent set of encoded versions of the media content is available. Within a period (610), material is arranged into Adaptation Sets (620). • Adaptation Set: an adaptation set (620) represents a set of interchangeable encoded versions of one or several media content components and contains a set of Representations (630). In some examples, the concept of Preselection (612) is used to enable a combination of different Adaptation Sets (630) into a single decoding instance and user experience. • Representation: a representation (630) describes a deliverable encoded version of one or several media content components and includes one or more media streams (e.g., one for each media content component in the multiplex). One example way to use the representation (630) is to apply different bit rates. Within a representation (630), the content may be divided in time into Segments (632, 634, 636) for proper accessibility and delivery. Docket No.: D23069WO01 • Segment: in some examples, a segment is a largest unit of data that can be retrieved with a single HTTP request. In order to access a segment, a URL is provided for each Segment. For segmented Representations, two types of Segments are used: o An Initialization Segment (632) contains static metadata for the Representation (630); o Media Segments (634) contain media samples and advance the timeline. • Self-initializing Segment: Some representations (630) may also be organized to have a single self-initializing Segment (636) which contains both initialization information and media data. [0067] FIG.7 is a block diagram illustrating example practical applications assigned to various hierarchical levels of the data model of FIG.6 according to some embodiments. In the example shown, the corresponding MPD has four periods (610). The periods (610) can be used for a first practical application (702) directed at splicing of arbitrary content, e.g., ad insertion, etc. Each period (610) includes a plurality of adaptation sets (620), which can be different types of multimedia. A corresponding second practical application (710) is directed at selection of components or tracks based on properties. In each adaptation set (620), we have different representations (630), each using a different respective bit rate. A corresponding third practical application (720) is directed at selecting and/or switching representations (630) based on communication-channel bandwidth. In each representation 630, there are multiple segments, e.g., including an initialization segment (632) and one or more media segments (634). A corresponding fourth practical application (730) provides a well-defined media format for the individual segments. Each segment (632, 634) contains a uniform resource locator (URL) address pointing to an mp4 bitstream. A corresponding fifth practical application (740) is directed at providing a media delivery format and data chunks with unique addresses and associated timing. [0068] FIG.8 is a pseudocode illustrating a common tag (800) using which an MPD is described in the XML format according to some examples. ISO Base Media File Format (BMFF) [0069] ISO BMFF uses “boxes” to structure different categories of data. Each box is led by 4 character type (4CC). Boxes can be nested and organized in hierarchical way, e.g., a box can be inside another box. The following piece of pseudocode provides an illustration of the box concept. Docket No.: D23069WO01 • “size” is an integer that specifies the number of bytes in this box, including all its fields and contained boxes; if “size” is 1 then the actual size is in the field “largesize”; if “size” is 0, then this box is the last one in the file, and its contents extend to the end of the file (normally only used for a Media Data Box); • “type” identifies the box type; standard boxes use a compact type, which is normally four printable characters, to permit ease of identification, and is shown so in the boxes below. User extensions use an extended type; in this case, the “type” field is set to ‘uuid’. [0070] Besides the above-illustrated “box,” a “FullBox,” with more information, can be defined as follows: Two different definitions are used to separate timed data and non-timed data. Docket No.: D23069WO01 • Track: image sequences are stored as tracks. An image sequence track is used when there is coding dependency between images or when the playback of the images is timed. Unlike in video tracks, the timing in the image sequence track is advisory. • Item: still images are stored as items. All image items are independently coded and do not depend on any other item for their decoding. Any number of image items can be included in the same file. [0071] Timed data in an ISO BMFF file are divided into samples which are described using a logical concept called “track.” A track box, identified by the 4CC ‘trak’, is stored within a movie box, ‘moov’, and contains number of boxes that provide information on where to find the samples in the file and how to interpret them. The samples themselves are stored in a media data box identified by the 4CC ‘mdat’. Each track contains at least one sample entry within a sample description box, ‘stsd’. A sample entry describes the coding and encapsulation format used in the samples of the track. Sample entries are also identified by a 4CC and may contain more boxes for further configurations. In some examples, non-timed data are stored in the meta box, ‘meta’, which is represented by one sample described by a logical concept called an item. Non- timed sample data may either be stored in the media data box or in an item data box, ‘idat’, that is stored in the meta box. [0072] The ISO BMFF defines the top-level boxes as follows: It is noted that one sample entry type defined by the ISO BMFF specification is the restricted video sample entry identified by the 4CC ‘resv’. This sample entry is used when the decoded sample of a track needs postprocessing for rendering or other purposes. The postprocessing operations are identified by a scheme type signaled within the restricted sample entry itself. For the case of signaling of supplemental enhancement information (SEI), a file author can list occurring SEI message IDs and classify them into two categories: those that are deemed required by the file author for correct playback, and others. The occurrence of either type of SEI Docket No.: D23069WO01 messages can be signalled in the SEI Information box using ‘seii’ under the scheme information box (‘schi’). The scheme for signalling of SEI uses ‘aSEI’ as scheme_type. [0073] FIG.9 shows Table 1 that provides an example of a restricted video sample entry ‘resv’. The restricted sample entry is useful, e.g., for an MPI video, because MPI uses postprocessing based on the MPI information SEI message and camera information. [0074] FIG.10 is a block diagram illustrating an overall structure of an ISO BMFF file (1000) for timed data according to some examples. The legend shown in FIG.10 provides more details describing the various boxes of the file (1000). Several boxes deemed to be particularly relevant to example embodiments disclosed herein are indicated by the arrows. [0075] FIG.11 is a block diagram illustrating an overall segment structure of fragmented MP4 used in DASH according to some examples. The shown example contains an initialization segment (1102) and a media segment (1104). The initialization segment (1102) contains a ‘moov’ box. In some examples, the ‘moov’ box is a mandatory box, and there should be exactly one box. This box contains other boxes that have detailed metadata information about the representation. The media segment (1104) contains a ‘moof’ box and an ‘mdat’ box. In some examples, the ‘moov’ box contains the metadata for the whole movie, but the ‘moof’ box is used to allow the tracks (e.g., audio, video, and text) to be broken up into fragments and contains the corresponding metadata for each track. There is one ‘moof’ box per fragment. The media data box ‘mdat’ holds the actual media samples for a representation. Each Media Segment contains one or more whole self-contained movie fragments. A whole, self-contained movie fragment is a movie fragment ('moof') box and a media data ('mdat') box that contains all the media samples that do not use external data references referenced by the track runs in the movie fragment box. Each 'moof' box contains at least one track fragment. [0076] FIG.12 is a block diagram illustrating an overall structure of a DASH representation (1200) according to some examples. In the example shown, the representation (1200) includes an initialization segment (1202) and a plurality of media segments (1204). The initialization segment (1202) contains static metadata for the representation (1200). The media segments (1204) contain media samples and advance the timeline. As already indicated above, ISO BMFF supports fragmented design, such as that of representation (1200). MPI Transmission Using DASH and ISO BMFF Docket No.: D23069WO01 [0077] FIG.13 is a block diagram illustrating a communication system (1300) configured to support MPI transmissions using DASH and ISO BMFF according to some embodiments. The communication system (1300) includes a server (1302) and a client (1306) connected via a communication channel (1304). The server (1302) includes a data storage device having stored therein one or more MPI containers (1400). In operation, the client (1306) requests and receives, via the communication channel (1304), pertinent portions of the data stored in the MPI containers (1400) as described in more detail below. [0078] FIG.14 is a block diagram illustrating a data structure of the MPI container (1400) used in the communication system (1300) according to some embodiments. The MPI container (1400) includes a plurality of segments logically organized (e.g., using appropriate indexing) to form three sets of planes of a 3D space whose dimensions are Adaptation Set, Representation, and Segment Time. There are two types of segments: initialization segments (632) and media segments (634). An end user at the client (1306) will typically request one initialization segment (632) per period (610). In one example, the initialization segment (632) contains camera information and metadata related to the media segment component (e.g., video, audio, subtitle, etc.) so that the end user knows which view(s) to use. An individual media segment (634) contains the audio and/or video content of a different respective view at a different respective bit rate in the corresponding period (610). The end user, through the client (1306), will typically request only the needed media segment(s) (634). [0079] Each Adaptation Set includes the Representation (630) indexed to logically be in a plane that is orthogonal to the Adaptation Set axis of the logical 3D space of the MPI container (1400). In one example, each Adaptation Set has MPIs from one camera view. If we have fifty camera views, then the MPI container (1400) has fifty corresponding Adaptation Sets. The end user selects a specific Adaptation Set by choosing which camera MPI to render. In various examples, the end user may request segments from one Adaptation Set for single-view MPI rendering or segments from multiple Adaptation Sets for multi-MPI rendering, depending on the rendering algorithm and other relevant conditions. [0080] Each Representation includes the segments (632, 634) indexed to logically be in a plane that is orthogonal to the Representation axis of the logical 3D space of the MPI container (1400). In one example, each Representation is associated with one respective encoded MPI bit rate. Different Representations typically correspond to different respective bit rates. The segments (632, 634) of various Representations are chosen based on the current condition of the communication channel (1304) between the server (1302) and the client (1306). Docket No.: D23069WO01 [0081] Each segment-time plane has the segments (634) indexed to logically be in a plane that is orthogonal to the Segment Time axis of the logical 3D space of the MPI container (1400). In one example, each segment-time plane is associated with a different respective playtime within a video. Different segment-time planes typically correspond to different respective playtimes. In the view presented in FIG.14, each of the segments (634) is shown as a box whose location in the 3D space of the MPI container (1400) reflects the segment’s playtime, view, and bit rate. [0082] In some examples, to enable the end user to select, via the client (1306), which media segment(s) (634) to download, the metadata transmitted through the communication channel (1304) include: • An MPD (1312): the MPD (1312) provides information about the MPI container (1400) and the URLs for different segments corresponding to different Adaptation Set, representation, and segment-time planes. The MPD (1312) enables the end user to retrieve sufficient MPI scene information and determine which media bitstream(s) to download. • An initialization segment (1314): an initialization segment (1314) is sent for each period (610) and provides information about the number of views, the intrinsic/extrinsic camera parameters, and indicators for post-decoder operations along the time dimension. Note that different MPI scenes may have different respective numbers of cameras and different respective camera-pose setups. The transmitted initialization segment (1314) contains those time-dependent pieces of information. [0083] At the client (1306), three inputs are used to determine which media segment(s) to request from the server (1302): • Novel viewing pose (1336): o In some examples, from the user input device (eye tracking, touch screen, sensors, etc.), the client (1306) infers the users’ new selected viewing pose (1336). o From the initialization segment (1314), the client (1306) has information about the selection of camera views. o Based on the above-indicated information, the client (1306) operates to determine from which view(s) (Adaptation Set(s)) of the MPI container (1400) to request segments. o Note that depending on the rendering algorithm, the segments can be requested from one Adaptation Set or multiple Adaptation Sets of the MPI container (1400). Docket No.: D23069WO01 o In some examples, a specialized algorithm configured to provide smooth playback transitions when the end-user requests a relatively large change in the novel viewing pose (1336) can be used to request a larger number of neighboring views, for example, at a lower bit rate due to bandwidth considerations. • Network condition (1332): According to the network condition experienced by the communication channel (1304), such as bandwidth, packet loss, and packet delay, the client (1306) operates to determine from which Representation of the MPI container (1400) to request segments. • Current time (1334): According to the current POC (Picture Order Count), the client (1306) “knows” how many frames are left in the current segment and which segment needs to be requested next from the server (1302) for continuous playback. Using the above-indicated inputs (1332, 1334, 1336), the client (1306) operates to generate a corresponding request (1326) for the desired media segments (634) needed for the end user continue seamless MPI playback. The request (1326) is then communicated, e.g., via an http GET message (1316) to the server (1302). In response to the message (1316), the server (1302) transmits MPI bitstreams (1318) carrying the desired media segments (634) for being rendered (1338) at the client (1306). [0084] In various additional embodiments, the MPI container (1400) can be further extended to include additional dimensions, such as resolution, codec type, frame rate, etc. In some examples, the bit rate dimension may be collapsed down to a single bit-rate value. MPD [0085] Some embodiments of the MPD (1312) are already described above in reference to FIGS.5-8. Additional embodiments of the MPD (1312) are described below in reference to FIGS.15-18. Two example use cases are covered by those additional embodiments: (i) a single- view MPI use case and (ii) a multi-view MPI use case. [0086] FIG.15 shows Table 2 that provides an example of the MPD (1312) for single-view MPI. The MPI-related information can be extracted from the Initialization Segment (632) (“init.mp4"), e.g., as shown in FIG.12, or from the ‘moov’ when there is only one segment in the Representation dimension of the MPI container (1400). In the example shown, the value of @codecs is set to “resv.mpiv.XXXX,” where XXXX corresponds to the 4CC of the video codec from the original_format field in RestrictedSchemeInfoBox of Sample Entry (e.g., 'avc1' or 'hvc1'). This feature can be used to inform the client player relatively quickly that the video is in Docket No.: D23069WO01 the MPI packing format so that player can transition into a corresponding appropriate configuration. In another example, the value of @codec is set to “XXXX,” and parsing of the Initialization Segment (632) is performed. [0087] In some examples, we would like to present the user with free-view rendering experience in a multi-view MPI system. Toward that goal, the client (1306) needs to “know” the multi-camera (reference views) setup from the MPD. Then, the client (1306) can request the corresponding reference view MPI video bitstream(s) based on the input (1336) as described above. Because all needed information for free-view rendering is captured in at least some embodiments of the above-described ISO BMFF design, the following basic format can be used for encapsulation and signaling in DASH: - Each MPI reference view video is represented in the DASH MPD as a separate Adaptation Set. - Each Adaptation Set contains a corresponding Initialization Segment (632). [0088] When the client (1306) accepts the MPD (1312), the client (1306) can construct the cameras setup based on the metadata information from the Initialization Segment (1314) from all the views. In some examples, the client (1306) can use ‘mvcg’ to identify prefetched neighbors and request low resolution/low bitrate video to shorten the delay for the free-view rendering (1338). In some examples, the client (1306) can also be configured to use the element Viewpoint in DASH for this purpose, e.g., as follows: [0089] FIG.16 shows Table 3 that provides an example of the MPD (1312) for multi-view MPI. For brevity, some portions of the MPD (1312) are not explicitly shown. Based on the description provided herein, a person of ordinary skill in the pertinent art will be able to add those portions back in without any undue experimentation. In the example shown, view 0, view 1, etc., are put into different respective adaptation sets, identified by the corresponding different Viewpoint values, containing different respective bitrates. Each adaptation set contains a corresponding Initialization Segment (632). Docket No.: D23069WO01 [0090] FIG.17 shows Table 4 that provides another example of the MPD (1312) for multi- view MPI. The example of FIG.17 can be used, e.g., when the user does not want to have any interaction for free-view rendering, and as such is directed to watch a director-recommended fly- through view. In some cases, this example can be used to also provide the initial view. [0091] FIG.18 shows Table 5 that provides yet another example of the MPD (1312) for multi-view MPI. In this example, one Adaptation Set is included, which only has only one Representation. That Representation only includes one Initialization Segment (632), but no Media Segments (634). The Initialization Segment (632) contains the Multiview global information. In this manner, the client (1306) only needs to download one Initialization Segment (632) and to get the global camera setup. Initialization Segment Bitstream [0092] To enable novel view rendering using an MPI scene representation, besides having the MPI side information, also needed is the camera information, including both intrinsic and extrinsic camera parameters. Because MPI uses postprocessing based on the MPI information SEI message and camera information, which falls into the post-decoder input needed for handling the media, some embodiments disclosed herein are directed to the storage of video component tracks. In some examples, such storage utilizes an existing capability of the ISO BMFF and a corresponding ISO/IEC 14496-15 mechanism. This mechanism enables players to inspect a file to determine the capabilities for rendering a bitstream and, as such, stops legacy players from decoding and rendering files for which such players lack appropriate processing capabilities. Therefore, important pieces of information contained in the initialization segment (632) are: (i) the camera pose, including the intrinsic and extrinsic camera model; and (ii) an indicator showing a need for post-decoder operation using SEI. In an example scenario considered below, we have N cameras. Herein below, we first discuss these pieces of information in more detail and then discuss how to organize these pieces of information in different Adaptive Sets. [0093] For the carriage of MPI coding data in ISO BMFF, various embodiments described herein blow address the following issues: - View identifier (relates to the multi-camera setup); - Carriage of intrinsic camera information; - Carriage of extrinsic camera information; and - Signalling of post-decoder operations using SEI. Docket No.: D23069WO01 As used in this specification, about the term “camera” should be interpreted to encompass both a true camera and a virtual camera. In some examples, for a virtual camera, the MPI data can be rendered from camera interpolation, 3D model rendering (e.g., Neural Radiance Fields (NeRF)), or using another suitable approach. Some embodiments leverage existing ISO BMFF capabilities (e.g., described in ISO/IEC 14496-15), whereas some other embodiments may employ a new box, a new scheme, and/or a new type geared specifically to enhanced or new functions implemented for MPI video streaming. [0094] FIGS.19A-19F provide definitions related to a view identifier used in the communication system (1300) according to some embodiments. The view identifier is represented by the “vwid” box, which may be defined as indicated in FIGS.19A and 19B. FIG. 19C provides an example syntax of view_id, which is generally suitable for MPI transmission purposes in the communication system (1300). However, the semantics of view_id shown in FIG.19C indicates the value of the view_id syntax element in the NAL unit header extension, which may not be present for a single layer codec. To address this problem and also the related problems of unusable syntax in some configurations, FIGS.19D-19F define an additional syntax provisionally labeled as version=1 and add further sample entries to allow inclusion of HEVC and VVC use cases. In some applications, when backward compatibility is not required or cannot be used, some alternative embodiments may use a new defined box (e.g., ‘mvid’) to substitute the ‘vwid’ version=1. [0095] FIGS.20A and 20B provide example definitions related to the intrinsic camera parameters used in the communication system (1300) according to some embodiments. In the example shown, the intrinsic camera parameters box ‘icam’ is compatible with ISO/IEC 14496- 15. [0096] FIGS.21A and 21B provide example definitions related to the extrinsic camera parameters used in the communication system (1300) according to some embodiments. In the example shown, the intrinsic camera parameters box ‘ecam’ is compatible with ISO/IEC 14496- 15. For MPI coded video, we can reuse ‘icam’ and ‘ecam’ boxes by extending container Sample Entry to include HEVC and VVC, etc., for example as follows: Docket No.: D23069WO01 Note that, in some examples, the value of syntax ‘view_id’ in box ‘vwid’, the value of syntax ‘ref_view_id’ in boxes ‘ecam’ and ‘icam’ is set equal to mpi_view_id in MPI information SEI message for the same camera view. [0097] FIG.22 provides example definitions related to a restricted scheme information used in the communication system (1300) according to some embodiments. In some examples, the initialization segment (632) incorporates metadata related to the video components. Such metadata mainly describes MPI packing information and is placed under the box ‘resv’. MPI coded video invokes a post-decoder mechanism. Therefore, an MPI video component track can be represented in the file as restricted video and can use a generic restricted sample entry ‘resv’. The restricted scheme information box ‘rinf’ can contain OriginalFormatBox (frma) to document the original sample entry type and a SchemeTypeBox (schm). In some examples, a SchemeInformationBox (schi) may also be used, depending on the restriction scheme. [0098] In some examples, two methods (2400, 2600) for MPI video arrangement are implemented in the box ‘resv’. These two methods are described in more detail below in reference to FIGS.23-26. [0099] In some examples, the MPI information SEI packing arrangements support both spatially packed constituent frames for texture packed image and alpha map packed image or temporally interleaved texture packed image and alpha map packed image. For the temporally interleaved case, one can put the texture packed image and the alpha map packed image in a single track or different tracks. In some examples, the system (1300) is configured to signal the MPI related SEI using the mechanisms described in ISO/IEC 14496-15. For example, the SchemeType ‘aSEI’ can be used. Then, one can use the SEI information box ‘seii’ under ‘schi’. [00100] FIGS.23A and 23B provide example definitions related the SEI information box ‘seii’ used in the communication system (1300) according to some embodiments. In ‘seii’, numRequiredSEIs is larger than 0. It sets one requiredSEI_ID to be the value of the MPI information SEI message “payloadType”. Docket No.: D23069WO01 [00101] FIG.24 shows Table 6 that provides an example of the method (2400) for MPI video arrangement with an MPI-related SEI message used in the communication system (1300) according to some embodiments. For illustration purposes and without any implied limitations, the method (2400) is expressed in terms of the Extensible Markup Language (XML). [00102] FIG.25 provides example definitions related to MPI video arrangement (MPIVideoBox) used in the communication system (1300) according to some embodiments. These definitions are compatible with the design for 8.1.1 scheme for stereoscopic video arrangement of ISO/IEC 14496-12. It is noted that these definitions are directed at keeping MPIVideoBox relatively streamlined, e.g., to contain a streamlined set of information. In the example shown, we only put mpi_indication_type for highest level information. In various examples, we may add the number of MPI layers and whether the MPI has constant depth distant among layers or not in the MPIVideoBox. For other details, we may use the MPI information SEI message. It is further noted that, in some examples, we can have two ‘resv’ entries under ‘stsd’, e.g., one for the MPI information SEI and another one for ‘mpiv’. [00103] FIG.26 shows Table 7 that provides an example of the method (2600) for MPI video arrangement using MPIVideoBox used in the communication system (1300) according to some embodiments. For illustration purposes and without any implied limitations, the method (2400) is expressed in terms of the XML. [00104] Each of AVC, HEVC, VSEI provides for the use of a multiview acquisition information SEI (MAI SEI) message, which contains the intrinsic parameter and the extrinsic parameter for perspective projection. In some examples, the system (1300) is configured to reuse the MAI SEI message to describe the camera information as well. Accordingly, the following modifications to Table 6 (FIG.24) can be implemented: remove the ‘icam’ and ‘ecam’ entry while changing the value of numRequiredSEIs to 2 and adding requiredSEI_ID of MAE SEI payload type. When comparing the above two solutions, the use of the ‘icam’ and ‘ecam’ boxes may prove to be more versatile, e.g., because those boxes are at a high level, and bitstreams do not need to be parsed to get the camera information. [00105] The following description addresses organization of information in the initialization segment (632) according to some examples. More specifically, the following several methods to organize the aforementioned information in different adaptation sets can be used. • Single-View (local) Info: the camera information and MPI-related information for each view is stored in its corresponding view’s initialization segment (632). In this case, in each Docket No.: D23069WO01 period (610), the client (1306) operates to request N initialization segments from the server (1302) to obtain all relevant cameras’ information. • Multi-View (global) Info: the camera information and MPI-related information for all views are stored in one view’s initialization segment (632). In this case, the client (1306) operates to request one initialization segment for each period to obtain those N cameras’ info. In the case of Single-View (local) Info, each view contains its own ‘icam’ and ‘ecam’. The client (1306) operates to request all initialization segments from all views of the MPI container (1400). In the case of Multi-View (global) Info, from ISO-BMFF, the “quantity” of ‘icam’ and ‘ecam’ can be 0 or more. Therefore, we can carry multiple cameras information using multiple ‘icam’ and ‘ecam’ boxes. For example, in the box ‘vwid’, we specify syntax num_views and view_id information. In box ‘ecam’ and box ‘icam’, we specify syntax ‘ref_view_id’ to map ‘ecam’ and ‘icam’ to corresponding syntax ‘view_id’ in box ‘vwid’. All their values are equal to mpi_view_id in MPI information SEI message for the same camera view. [00106] FIG.27 shows Table 8 that provides an example used to specify a multi-view camera setup in the communication system (1300) according to some embodiments. For illustration purposes and without any implied limitations, the example is expressed in terms of the XML. Media Segment Bitstream [00107] In various examples, the media segment (634) has packed therein one or more MPI bitstreams and the corresponding MPI SEI information. In cases of video and/or audio bitstreams, the media segment (634) contains video samples (bitstreams), audio samples (bitstreams), etc. The MPI related SEIs are carried in video bitstreams. In some examples, MPI SEI information provides certain MPI information, such as the number of MPI layers, the packing of texture and alpha data for the MPI layers, and the depth value for each layer. In addition, it specifies mpi_view_id corresponding to multi-camera setups. Pre-Fetch for Streaming [00108] In some cases, in order to reduce delays, a prefetch strategy is adopted for streaming. Therefore, it is important for implementing the prefetch strategy that the client (1306) takes actions to identify the views to be prefetched. ISO/IEC 14496-15 provides a multiview group box (‘mvcg’) contained in the multiview information box (‘mvci’, contained in ‘minf’). In some examples, the system (1300) can be configured to use ‘mvcg’ to group neighboring views. In one example, for each view, the system (1300) can indicate the current view with eight Docket No.: D23069WO01 neighboring views into one Multiview Group Box ‘mvcg’. The box ‘mvcg’ is contained in ‘mvci’ (Multiview Information box), and ‘mvci’ is contained in ‘minf’(Media Information Box). [00109] The box ‘mvcg’ was originally designed to specify a Multiview group for the views of the MVC and MVD stream(s) that are output. Target views can be indicated on the basis of track_id (e.g., entry_type equal to 0) within the Multiview Group box. When a multiview sample grouping is in use, and tiers cover more than one view or some tiers contain a temporal subset of the bitstream, it is recommended to use tier_id (i.e., entry_type equal to 1) within the Multiview Group box. Otherwise, it is recommended to use one of the view_id based indications (i.e., entry_type equal to 2 or 3). [00110] FIGS.28A-28C provide example definitions related pre-fetch operations used in the communication system (1300) according to some embodiments. In the example shown, we reuse ‘mvcg’ for the above-indicated purposes. To be more specific, we use entry_type equal to 2 to specify the syntax output_view_id. The output_view_id corresponds to one neighboring camera view_id of the current camera view. For example, to specify 8 neighboring views, we can set multiview_group_id to the current camera view_id and num_entries to 8. Storage Format [00111] In some examples, for the streaming (DASH) case, MPI for each view is stored in a different respective ISO BMFF file. FIG.29 is a block diagram (2900) illustrating encapsulation of multi-view MPI coded data in one ISO BMFF for storage according to some embodiments. In some examples, for storage purposes (as opposed to streaming), to encapsulate Multiview MPI coded data in a single ISO BMFF file, the communication system (1300) is configured to put each view MPI coded data into separate tracks (2902-2908), indicated by syntax track_ID under ‘tkhd’, contained in ‘trak’ which is contained in the ‘moov’ box. This approach allows significant flexibility in terms of rendering a novel view, both spatially and temporally. Example Hardware [00112] FIG.30 is a block diagram illustrating a computing device (3000) used in the communication system (1300) according to an embodiment. The device (3000) can be used, e.g., at the server (1302) or at the client (1306). The computing device (3000) comprises input/output (I/O) devices (3010), a processing engine (3020), and a memory (3030). The I/O devices (3010) may be used to enable the device (3000) to receive various input signals (3002) Docket No.: D23069WO01 and to output various output signals (3004). For example, the I/O devices (3010) may be operatively connected to send and receive signals via the communication channel (1304). [00113] The memory (3030) may have buffers to receive data. Once the data are received, the memory (3030) may provide parts of the data to the processing engine (3020) for processing therein. The processing engine (3020) includes a processor (3022) and a memory (3024). The memory (3024) may store therein program code, which when executed by the processor (3022) enables the processing engine (3020) to perform various data processing operations, including but not limited to at least some operations of the above-described MPI methods. [00114] According to an example embodiment disclosed above, e.g., in the summary section and/or in reference to any one or any combination of some or all of FIGS.1-30, provided is an apparatus for MPI video streaming, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: provide to a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, provide to the client device a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; receive, from the client device, a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and transmit to the client device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended parameter values. As used herein, the term “view” should be interpreted to encompass both a captured view and a novel view. In some examples, novel views can be generated for MPI streaming based on captured views, e.g., via a NeRF 3D model. One benefit of using novel view positions in at least some examples of MPI streaming is to make the camera pose more uniform, such as to facilitate more-efficient multi-MPI rendering at the decoder. [00115] According to another example embodiment disclosed above, e.g., in the summary section and/or in reference to any one or any combination of some or all of FIGS.1-30, provided is a server-implemented method for MPI video streaming comprising: providing to a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, providing to the client device a respective Docket No.: D23069WO01 initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; receiving, from the client device, a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and transmitting to the client device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. [00116] In some embodiments of the above server-implemented method, for the period, the storage container has a plurality of media segments logically organized in accordance with different views and further logically organized in accordance with one or more of different bit rates, different resolutions, different codec types, and different frame rates. [00117] In some embodiments of any of the above server-implemented methods, for each of the different captured views and each of the different bit rates, the storage container has a respective sequence of media segments corresponding to different respective video segment times. [00118] In some embodiments of any of the above server-implemented methods, the method further comprises switching from a first respective sequence of media segments to a different second respective sequence of media segments when the request indicates a change in the identified selection or a change of the recommended bit rate. [00119] In some embodiments of any of the above server-implemented methods, the transmitting includes: transmitting a first bitstream carrying a first respective sequence of media segments corresponding to a first one of the different captured views; and transmitting a second bitstream carrying a different second respective sequence of media segments corresponding to a second one of the different captured views, wherein the first bitstream and the second bitstream have respective media segments corresponding to a same one of the different respective video segment times. [00120] In some embodiments of any of the above server-implemented methods, the media presentation description includes one or more information sets selected from the group consisting of: a captured-views information set; a bit-rate information set; a video-resolution information set; a frame rate information set; an audio-language information set; and an Docket No.: D23069WO01 information set identifying one or more uniform resource locator addresses of initialization segments or media segments stored in the storage container. [00121] In some embodiments of any of the above server-implemented methods, the media presentation description includes information corresponding to two or more periods of the MPI streaming content. [00122] In some embodiments of any of the above server-implemented methods, the media presentation description includes information corresponding to two or more different captured views. [00123] In some embodiments of any of the above server-implemented methods, the respective initialization segment includes camera information and metadata related to one or more components of the media segments (e.g., video, audio, subtitle, etc.). [00124] In some embodiments of any of the above server-implemented methods, the respective initialization segment has a box format compatible with at least one of an MPEG DASH specification and an ISO BMFF specification and/or it is extensions, e.g., CMAF (using ISO BMFF for defining a container), etc. [00125] In some embodiments of any of the above server-implemented methods, the respective initialization segment contains multiview information or single view information (e.g., camera information or view rendered from some model (such as Neural radiance field (NeRF)). [00126] In some embodiments of any of the above server-implemented methods, the method further comprises transmitting a supplemental-enhancement-information (SEI) message or an MPI video box specifying a packing arrangement of constituent frames, the packing arrangement being selected from: a first packing arrangement including one or more texture-packed and alpha-map packed images; and a second packing arrangement including temporally interleaved texture-packed images and alpha-map packed images, either in a single track or in two or more different tracks. [00127] In some embodiments of any of the above server-implemented methods, the method further comprises, for the period, providing to the client device two or more of the respective initialization segments from the storage container, each of the two or more respective initialization segments corresponding to a different respective view of the MPI streaming content. Docket No.: D23069WO01 [00128] In some embodiments of any of the above server-implemented methods, the one or more bitstreams include: a sequence of the media segments containing at least video samples and audio samples; and supplemental-enhancement-information specifying one or more MPI parameters of the MPI streaming content. [00129] In some embodiments of any of the above server-implemented methods, the method further comprises providing a view identifier box including one or more of: indication of views included in a track or in a tier; indication of view order index for each listed view; minimum and maximum values of a temporal parameter for the track or for the tier; indication of reference views for decoding the views included in the track or tier; and for each of the included views, indication of presence of texture or depth in the track. [00130] In some embodiments of any of the above server-implemented methods, the method further comprises providing a view identifier box or another box including one or both of: indication of views included in a track or in a tier; and indication of view identifier for each listed view. [00131] In some embodiments of any of the above server-implemented methods, the method further comprises providing an MPI packing information box using a restricted sample entry defined in an ISO BMFF specification. [00132] In some embodiments of any of the above server-implemented methods, the method further comprises providing an MPI video box to indicate whether decoded frames contain a representation of two spatially packed frames or a representation of two temporally interleaved constituent frames forming an MPI representation. [00133] In some embodiments of any of the above server-implemented methods, the method further comprises providing a multiview group box indicating a current view and a fixed number of neighboring views. [00134] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor of a server device, cause the server device to perform operations comprising any one of the above server-implemented methods for MPI video streaming. [00135] According to yet another example embodiment disclosed above, e.g., in the summary section and/or in reference to any one or any combination of some or all of FIGS.1-30, provided is an apparatus for MPI video streaming, the apparatus comprising: at least one processor; and at Docket No.: D23069WO01 least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: receive at a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, receive at the client device, via the server device, a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; transmit to the server device a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and receive from the server device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. [00136] According to another example embodiment disclosed above, e.g., in the summary section and/or in reference to any one or any combination of some or all of FIGS.1-30, provided is a client-implemented method for MPI video streaming comprising: receiving at a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, receiving at the client device, via the server device, a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; transmitting to the server device a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and receiving from the server device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. [00137] In some embodiments of the above client-implemented method, for the period, the storage container has a plurality of media segments logically organized in accordance with different captured views and further logically organized in accordance with one or more of different bit rates, different resolutions, different codec types, and different frame rates. [00138] In some embodiments of any of the above client-implemented methods, for each of the different captured views and each of the different bit rates, the storage container has a respective sequence of media segments corresponding to different respective video segment times. Docket No.: D23069WO01 [00139] In some embodiments of any of the above client-implemented methods, the method further comprises requesting a change in the identified selection or a change of the recommended bit rate to cause the server device to switch from a first respective sequence of media segments to a different second respective sequence of media segments. [00140] In some embodiments of any of the above client-implemented methods, the method further comprises receiving a first bitstream carrying a first respective sequence of media segments corresponding to a first one of the different captured views; and receiving a second bitstream carrying a different second respective sequence of media segments corresponding to a second one of the different captured views, wherein the first bitstream and the second bitstream have respective media segments corresponding to a same one of the different respective video segment times. [00141] In some embodiments of any of the above client-implemented methods, the media presentation description includes one or more information sets selected from the group consisting of: a captured-views information set; a bit-rate information set; a video-resolution information set; a frame rate information set; an audio-language information set; and an information set identifying one or more uniform resource locator addresses of initialization segments or media segments stored in the storage container. [00142] In some embodiments of any of the above client-implemented methods, the media presentation description includes information corresponding to two or more periods of the MPI streaming content. [00143] In some embodiments of any of the above client-implemented methods, the media presentation description includes information corresponding to two or more different captured views. [00144] In some embodiments of any of the above client-implemented methods, the respective initialization segment includes camera information and metadata related to one or more components of the media segments. [00145] In some embodiments of any of the above client-implemented methods, the respective initialization segment has a box format compatible with at least one of an MPEG DASH specification, an ISO BMFF specification, and a CMAF specification. Docket No.: D23069WO01 [00146] In some embodiments of any of the above client-implemented methods, the respective initialization segment contains multiview information or single view information. [00147] In some embodiments of any of the above client-implemented methods, the method further comprises receiving a supplemental-enhancement-information (SEI) message or an MPI video box specifying a packing arrangement of constituent frames, the packing arrangement being selected from: a first packing arrangement including one or more texture-packed and alpha-map packed images; and a second packing arrangement including temporally interleaved texture-packed images and alpha-map packed images, either in a single track or in two or more different tracks. [00148] In some embodiments of any of the above client-implemented methods, the method further comprises, for the period, receiving via the server device two or more of the respective initialization segments from the storage container, each of the two or more respective initialization segments corresponding to a different respective view of the MPI streaming content. [00149] In some embodiments of any of the above client-implemented methods, the one or more bitstreams include: a sequence of the media segments containing at least video samples and audio samples; and supplemental-enhancement-information specifying one or more MPI parameters of the MPI streaming content. [00150] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor of a client device, cause the client device to perform operations comprising any one of the above client-implemented methods for MPI video streaming. [00151] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims. [00152] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples Docket No.: D23069WO01 provided would be apparent upon reading the above description. The scope should be determined, not with reference to the above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation. [00153] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary. [00154] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter. [00155] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims. [00156] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit. [00157] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program code recorded in tangible media, such as magnetic recording media, optical recording media, Docket No.: D23069WO01 solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non-transitory machine-readable storage medium including being loaded into and/or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits. [00158] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range. [00159] The use of figure numbers and/or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures. [00160] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence. [00161] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.” [00162] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objects so referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner. Docket No.: D23069WO01 [00163] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if” may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].” [00164] Also for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements. [00165] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard, and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard. [00166] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and/or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and/or custom, may also be included. Similarly, any switches shown in the figures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context. Docket No.: D23069WO01 [00167] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and/or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and/or digital hardware circuit(s) with software/firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device. [00168] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown. [00169] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and/or in reference to one or more drawings. “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.

Claims

Docket No.: D23069WO01 CLAIMS What is claimed is: 1. A method for multi-plane-image (MPI) video streaming, comprising: providing to a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, providing to the client device a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; receiving, from the client device, a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and transmitting to the client device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. 2. The method of claim 1, wherein, for the period, the storage container has a plurality of media segments logically organized in accordance with different views and further logically organized in accordance with one or more of different bit rates, different resolutions, different codec types, and different frame rates. 3. The method of claim 2, wherein, for each of the different views and each of the different bit rates, the storage container has a respective sequence of media segments corresponding to different respective video segment times. 4. The method of claim 3, further comprising switching from a first respective sequence of media segments to a different second respective sequence of media segments when the request indicates a change in the identified selection or a change of the respective recommended values. 5. The method of claim 3, wherein the transmitting includes: transmitting a first bitstream carrying a first respective sequence of media segments corresponding to a first one of the different views; and Docket No.: D23069WO01 transmitting a second bitstream carrying a different second respective sequence of media segments corresponding to a second one of the different views, wherein the first bitstream and the second bitstream have respective media segments corresponding to a same one of the different respective video segment times. 6. The method of claim 1, wherein the media presentation description includes one or more information sets selected from the group consisting of: a views information set; a bit-rate information set; a video-resolution information set; a frame rate information set; an audio-language information set; and an information set identifying one or more uniform resource locator addresses of initialization segments or media segments stored in the storage container. 7. The method of claim 1, wherein the media presentation description includes information corresponding to two or more periods of the MPI streaming content. 8. The method of claim 1, wherein the media presentation description includes information corresponding to two or more different views. 9. The method of claim 1, wherein the respective initialization segment includes camera information and metadata related to one or more components of the media segments. 10. The method of claim 1, wherein the respective initialization segment has a box format compatible with at least one of an MPEG DASH specification, an ISO BMFF specification, and a CMAF specification. 11. The method of claim 1, wherein the respective initialization segment contains multiview information or single view information. 12. The method of claim 1, further comprising transmitting a supplemental-enhancement- information (SEI) message or an MPI video box specifying a packing arrangement of constituent frames, the packing arrangement being selected from: Docket No.: D23069WO01 a first packing arrangement including one or more texture-packed and alpha-map packed images; and a second packing arrangement including temporally interleaved texture-packed images and alpha-map packed images, either in a single track or in two or more different tracks. 13. The method of claim 1, further comprising: for the period, providing to the client device two or more of the respective initialization segments from the storage container, each of the two or more respective initialization segments corresponding to a different respective view of the MPI streaming content. 14. The method of claim 1, wherein the one or more bitstreams include: a sequence of the media segments containing at least video samples and audio samples; and supplemental-enhancement-information specifying one or more MPI parameters of the MPI streaming content. 15. The method of claim 1, further comprising providing a view identifier box or another box including one or both of: indication of views included in a track or in a tier; and indication of view identifier for each listed view. 16. The method of claim 1, further comprising providing an MPI packing information box using a restricted sample entry defined in an ISO BMFF specification. 17. The method of claim 1, further comprising providing an MPI video box to indicate whether decoded frames contain a representation of two spatially packed frames or a representation of two temporally interleaved constituent frames forming an MPI representation. 18. The method of claim 1, further comprising providing a multiview group box indicating a current view and a fixed number of neighboring views. 19. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor of a server device, cause the server device to perform operations comprising any one of the methods of claim 1-18. Docket No.: D23069WO01 20. An apparatus for multi-plane-image (MPI) video streaming, the apparatus comprising: at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: provide to a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, provide to the client device a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; receive, from the client device, a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and transmit to the client device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. 21. A method for multi-plane-image (MPI) video streaming, comprising: receiving at a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, receiving at the client device, via the server device, a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; transmitting to the server device a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and receiving from the server device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values. 22. The method of claim 21, wherein, for the period, the storage container has a plurality of media segments logically organized in accordance with different views and further logically Docket No.: D23069WO01 organized in accordance with one or more of different bit rates, different resolutions, different codec types, and different frame rates. 23. The method of claim 22, wherein, for each of the different views and each of the different bit rates, the storage container has a respective sequence of media segments corresponding to different respective video segment times. 24. The method of claim 23, further comprising requesting a change in the identified selection or a change of the recommended bit rate to cause the server device to switch from a first respective sequence of media segments to a different second respective sequence of media segments. 25. The method of claim 23, further comprising: receiving a first bitstream carrying a first respective sequence of media segments corresponding to a first one of the different views; and receiving a second bitstream carrying a different second respective sequence of media segments corresponding to a second one of the different views, wherein the first bitstream and the second bitstream have respective media segments corresponding to a same one of the different respective video segment times. 26. The method of claim 21, wherein the media presentation description includes one or more information sets selected from the group consisting of: a views information set; a bit-rate information set; a video-resolution information set; a frame rate information set; an audio-language information set; and an information set identifying one or more uniform resource locator addresses of initialization segments or media segments stored in the storage container. 27. The method of claim 21, wherein the media presentation description includes information corresponding to two or more periods of the MPI streaming content. 28. The method of claim 21, wherein the media presentation description includes information corresponding to two or more different views. Docket No.: D23069WO01 29. The method of claim 21, wherein the respective initialization segment includes camera information and metadata related to one or more components of the media segments. 30. The method of claim 21, wherein the respective initialization segment has a box format compatible with at least one of an MPEG DASH specification, an ISO BMFF specification, and a CMAF specification. 31. The method of claim 21, wherein the respective initialization segment contains multiview information or single view information. 32. The method of claim 21, further comprising receiving a supplemental-enhancement- information (SEI) message or an MPI video box specifying a packing arrangement of constituent frames, the packing arrangement being selected from: a first packing arrangement including one or more texture-packed and alpha-map packed images; and a second packing arrangement including temporally interleaved texture-packed images and alpha-map packed images, either in a single track or in two or more different tracks. 33. The method of claim 21, further comprising: for the period, receiving via the server device two or more of the respective initialization segments from the storage container, each of the two or more respective initialization segments corresponding to a different respective view of the MPI streaming content. 34. The method of claim 21, wherein the one or more bitstreams include: a sequence of the media segments containing at least video samples and audio samples; and supplemental-enhancement-information specifying one or more MPI parameters of the MPI streaming content. 35. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor of a client device, cause the client device to perform operations comprising any one of the methods of claim 21-34. 36. An apparatus for multi-plane-image (MPI) video streaming, the apparatus comprising: Docket No.: D23069WO01 at least one processor; and at least one memory including program code; and wherein the at least one memory and the program code are configured to, with the at least one processor, cause the apparatus at least to: receive at a client device a media presentation description of an MPI streaming content stored in a storage container accessible via a server device; for a period, receive at the client device, via the server device, a respective initialization segment from the storage container, the respective initialization segment being configured to inform a selection, at the client device, of one or more views of the MPI streaming content for which to request media segments for rendering; transmit to the server device a request identifying the selection and indicating a respective recommended value of at least one parameter selected from the group consisting of a bit rate, a resolution, a codec type, and a frame rate; and receive from the server device one or more bitstreams carrying the media segments selected in the storage container based on the identified selection and further based on one or more of the respective recommended values.
EP24739985.0A 2023-06-27 2024-06-24 Architecture for interactive multi-view multiplane-imaging video streaming Pending EP4736460A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202363510571P 2023-06-27 2023-06-27
PCT/US2024/035225 WO2025006378A1 (en) 2023-06-27 2024-06-24 Architecture for interactive multi-view multiplane-imaging video streaming

Publications (1)

Publication Number Publication Date
EP4736460A1 true EP4736460A1 (en) 2026-05-06

Family

ID=91853612

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24739985.0A Pending EP4736460A1 (en) 2023-06-27 2024-06-24 Architecture for interactive multi-view multiplane-imaging video streaming

Country Status (3)

Country Link
EP (1) EP4736460A1 (en)
CN (1) CN121569490A (en)
WO (1) WO2025006378A1 (en)

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9596447B2 (en) * 2010-07-21 2017-03-14 Qualcomm Incorporated Providing frame packing type information for video coding
GB2534136A (en) * 2015-01-12 2016-07-20 Nokia Technologies Oy An apparatus, a method and a computer program for video coding and decoding
CN111869221B (en) * 2018-04-05 2021-07-20 华为技术有限公司 Valid associations between DASH objects

Also Published As

Publication number Publication date
WO2025006378A1 (en) 2025-01-02
CN121569490A (en) 2026-02-24

Similar Documents

Publication Publication Date Title
CN110870321B (en) Region-wise packaging, content coverage, and signaling frame packaging for media content
CN115211131B (en) Apparatus, method and computer program for omni-directional video
KR102037009B1 (en) A method, device, and computer program for obtaining media data and metadata from an encapsulated bit-stream in which an operation point descriptor can be set dynamically
CN110431850B (en) Signaling important video information in network video streaming using MIME type parameters
KR102559862B1 (en) Methods, devices, and computer programs for media content transmission
CN109076229B (en) Areas of most interest in pictures
US11094130B2 (en) Method, an apparatus and a computer program product for video encoding and video decoding
US11477489B2 (en) Apparatus, a method and a computer program for video coding and decoding
JP7035088B2 (en) High level signaling for fisheye video data
CN111034203A (en) Processing omnidirectional media with dynamic zone-by-zone encapsulation
CA3169708A1 (en) Three-dimensional content processing methods and apparatus
CN114930869B (en) Method, apparatus and computer program product for video encoding and video decoding
GB2608399A (en) Method, device, and computer program for dynamically encapsulating media content data
US12587719B2 (en) Method, device, and computer program for dynamically encapsulating media content data
CN110870323A (en) Processing media data using omnidirectional media format
EP4736460A1 (en) Architecture for interactive multi-view multiplane-imaging video streaming
WO2019014216A1 (en) Enhanced region-wise packing and viewport independent hevc media profile
EP4736458A1 (en) Multi-view multiplane-imaging video streaming
CN117581551A (en) Method, device and computer program for dynamically encapsulating media content data
HK40016389A (en) Processing media data using an omnidirectional media format
HK40016389B (en) Processing media data using an omnidirectional media format
HK40009749B (en) Signaling important video information in network video streaming using mime type parameters
HK40009749A (en) Signaling important video information in network video streaming using mime type parameters

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE