WO2022257518A1 - 沉浸媒体的数据处理方法、装置、相关设备及存储介质 - Google Patents

沉浸媒体的数据处理方法、装置、相关设备及存储介质 Download PDF

Info

Publication number
WO2022257518A1
WO2022257518A1 PCT/CN2022/080257 CN2022080257W WO2022257518A1 WO 2022257518 A1 WO2022257518 A1 WO 2022257518A1 CN 2022080257 W CN2022080257 W CN 2022080257W WO 2022257518 A1 WO2022257518 A1 WO 2022257518A1
Authority
WO
WIPO (PCT)
Prior art keywords
camera
track
image
free
field
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2022/080257
Other languages
English (en)
French (fr)
Inventor
胡颖
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Tencent Technology Shenzhen Co Ltd
Original Assignee
Tencent Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tencent Technology Shenzhen Co Ltd filed Critical Tencent Technology Shenzhen Co Ltd
Priority to US17/993,852 priority Critical patent/US12395615B2/en
Publication of WO2022257518A1 publication Critical patent/WO2022257518A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N23/00Cameras or camera modules comprising electronic image sensors; Control thereof
    • H04N23/90Arrangement of cameras or camera modules, e.g. multiple cameras in TV studios or sports stadiums
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof
    • H04N13/10Processing, recording or transmission of stereoscopic or multi-view image signals
    • H04N13/106Processing image signals
    • H04N13/172Processing image signals image signals comprising non-image signal components, e.g. headers or format information
    • H04N13/178Metadata, e.g. disparity information
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof
    • H04N13/10Processing, recording or transmission of stereoscopic or multi-view image signals
    • H04N13/106Processing image signals
    • H04N13/111Transformation of image signals corresponding to virtual viewpoints, e.g. spatial image interpolation
    • H04N13/117Transformation of image signals corresponding to virtual viewpoints, e.g. spatial image interpolation the virtual viewpoint locations being selected by the viewers or determined by viewer tracking
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof
    • H04N13/10Processing, recording or transmission of stereoscopic or multi-view image signals
    • H04N13/106Processing image signals
    • H04N13/161Encoding, multiplexing or demultiplexing different image signal components
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof
    • H04N13/10Processing, recording or transmission of stereoscopic or multi-view image signals
    • H04N13/194Transmission of image signals
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof
    • H04N13/20Image signal generators
    • H04N13/204Image signal generators using stereoscopic image cameras
    • H04N13/243Image signal generators using stereoscopic image cameras using three or more two-dimensional [2D] image sensors
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N13/00Stereoscopic video systems; Multi-view video systems; Details thereof
    • H04N13/30Image reproducers
    • H04N13/366Image reproducers using viewer tracking
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/176Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/597Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding specially adapted for multi-view video sequence encoding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • H04N19/61Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding in combination with predictive coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/80Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation
    • H04N19/82Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation involving filtering within a prediction loop
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/90Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using coding techniques not provided for in groups H04N19/10-H04N19/85, e.g. fractals
    • H04N19/91Entropy coding, e.g. variable length coding [VLC] or arithmetic coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N23/00Cameras or camera modules comprising electronic image sensors; Control thereof
    • H04N23/60Control of cameras or camera modules
    • H04N23/698Control of cameras or camera modules for achieving an enlarged field of view, e.g. panoramic image capture

Definitions

  • This application relates to the field of audio and video, in particular to media data processing.
  • Immersive media refers to media content that can bring consumers an immersive experience.
  • Immersive media can also be called free-view video.
  • Free-view video is usually a camera array that shoots the same 3D scene from multiple angles to obtain different perspectives.
  • Depth images and/or texture images of these depth images and/or texture images constitute a free-view video.
  • the content consumption device can choose to decode certain images for consumption according to the current location of the user and the camera angle of view of each image source.
  • a large-scale atlas information data box is generally used to indicate the parameter information related to free-view video (such as the depth map collected by the camera and the resolution width and height of the texture map, and the camera identification corresponding to each view. symbols, etc.), omitting the rest of the gallery information in the gallery track.
  • Embodiments of the present application provide a data processing method, device, device, and storage medium for immersive media, which can encapsulate images collected by cameras with different viewing angles of an immersive media into multiple different tracks, and use the free viewing angle corresponding to each track
  • the information data box indicates the viewing angle information of the image source camera in each track, so that the content consumption device can select an appropriate image for decoding and consumption according to the viewing angle information in each track and the user's current location.
  • the embodiment of the present application provides a data processing method for immersive media.
  • the immersive media is composed of images taken by N cameras at different viewing angles.
  • the immersive media is encapsulated into M tracks, and one track encapsulates Images from at least one camera, N and M are both integers greater than 1, data processing methods include:
  • the free viewing angle information data box includes the viewing angle information corresponding to the i-th track, and i is an integer greater than or equal to 1 and less than or equal to M;
  • the image encapsulated in the ith track is decoded according to the viewing angle information in the free viewing angle information data box.
  • the embodiment of this application provides another immersive media processing method, including:
  • the immersive media is composed of images taken by N cameras at different viewing angles, one track encapsulates images from at least one camera, and the M tracks belong to the same track group , N and M are both integers greater than or equal to 1;
  • the free view angle information data box includes the view angle information corresponding to the i-th track; 1 ⁇ i ⁇ M , w ⁇ 1.
  • an embodiment of the present application provides a data processing device for immersive media, including:
  • the obtaining unit is used to obtain the free view angle information data box corresponding to the i-th track of the immersive media, the free view angle information data box includes the view angle information corresponding to the i-th track, and the immersive media is composed of N Composed of images taken by several cameras, the immersive media is packaged into M tracks, and images from at least one camera are packaged in one track, N and M are both integers greater than 1; i is greater than or equal to 1 and less than or equal to M an integer of
  • a decoding unit configured to decode the image encapsulated in the i-th track according to the viewing angle information in the free viewing angle information data box.
  • the embodiment of the present application provides another data processing device for immersive media, including:
  • An encapsulation unit configured to encapsulate immersive media into M tracks, the immersive media is composed of images taken by N cameras at different viewing angles, one track encapsulates images from at least one camera, and the M tracks Belong to the same orbital group, N and M are both integers greater than or equal to 1;
  • a generating unit configured to generate a free viewing angle information data box corresponding to the i-th track according to the encapsulation process of the w images in the i-th track, the free viewing angle information data box including the viewing angle information corresponding to the i-th track; 1 ⁇ i ⁇ M, w ⁇ 1.
  • the embodiment of the present application provides a content consumption device, including:
  • processor adapted to implement one or more computer programs
  • a computer storage medium storing one or more computer programs adapted to be loaded and executed by a processor:
  • the free angle of view information data box corresponding to the i-th track of the immersive media includes the angle of view information corresponding to the i-th track;
  • the immersive media is taken by N cameras at different angles of view Image composition, the immersive media is packaged into M tracks, and images from at least one camera are packaged in one track, N and M are both integers greater than 1, and i is an integer greater than or equal to 1 and less than or equal to M;
  • the image encapsulated in the ith track is decoded according to the viewing angle information in the free viewing angle information data box.
  • an embodiment of the present application provides a content production device, including:
  • processor adapted to implement one or more computer programs
  • a computer storage medium storing one or more computer programs adapted to be loaded and executed by a processor:
  • the immersive media is composed of images taken by N cameras at different viewing angles, one track encapsulates images from at least one camera, and the M tracks belong to
  • N and M are integers greater than or equal to 1;
  • the free view angle information data box includes the view angle information corresponding to the i-th track; 1 ⁇ i ⁇ M , w ⁇ 1.
  • an embodiment of the present application provides a storage medium, where the storage medium is used to store a computer program, and the computer program is used to execute the method in the above aspect.
  • an embodiment of the present application provides a computer program product including instructions, which, when run on a computer, cause the computer to execute the method in the above aspect.
  • an immersive medium is packaged into M tracks.
  • the immersive medium is composed of images taken by N cameras at different viewing angles.
  • One track may include images from at least one camera.
  • the M tracks are Belonging to the same track group, a scene in which an immersive video is encapsulated into multiple tracks is realized; in addition, the content production device generates a free-view information data box for each track, which is indicated by the free-view information data box corresponding to the i-th track.
  • the viewing angle information corresponding to the i-th track such as the specific viewing angle position of the camera, and then when the content consumption device decodes the image in the track according to the viewing angle information corresponding to the i-th track, it can ensure that the decoded video is closer to the user's location Match to improve the rendering effect of immersive video.
  • Fig. 1a is a schematic diagram of a user consuming 3DoF immersive media provided by an embodiment of the present application
  • Fig. 1b is a schematic diagram of a user consuming 3DoF+ immersive media provided by an embodiment of the present application
  • Fig. 1c is a schematic diagram of a user consuming 6DoF immersive video provided by an embodiment of the present application
  • Fig. 2a is an architectural diagram of an immersive media system provided by an embodiment of the present application.
  • Fig. 2b is a schematic diagram of an immersive media transmission solution provided by an embodiment of the present application.
  • FIG. 3a is a basic block diagram of a video encoding provided by an embodiment of the present application.
  • Fig. 3b is a schematic diagram of an input image division adopted in the embodiment of the present application.
  • FIG. 4 is a schematic flowchart of a data processing method for immersive media provided by an embodiment of the present application
  • FIG. 5 is a schematic flowchart of another data processing method for immersive media provided by an embodiment of the present application.
  • FIG. 6 is a schematic structural diagram of a data processing device for immersive media provided by an embodiment of the present application.
  • FIG. 7 is a schematic structural diagram of another data processing device for immersive media provided by an embodiment of the present application.
  • FIG. 8 is a schematic structural diagram of a content consumption device provided by an embodiment of the present application.
  • Fig. 9 is a schematic structural diagram of a content creation device provided by an embodiment of the present application.
  • the embodiment of the present application relates to the data processing technology of immersive media.
  • the so-called immersive media refers to media files that can provide immersive media content, so that users immersed in the media content can obtain visual, auditory and other sensory experiences in the real world.
  • immersive media can be three degrees of freedom (3Degree of Freedom, 3DoF) immersive media, 3DoF+ immersive media or six degrees of freedom (6Degree of Freedom, 6DoF) immersive media.
  • FIG. 1a it is a schematic diagram of a user consuming 3DoF immersive media provided by the embodiment of this application.
  • the 3DoF immersive media shown in Figure 1a means that the user is fixed at the center point of a three-dimensional space, and the user's head is along the X axis, Y axis Axis and Z-axis rotation to view the screen provided by the media content of immersive media.
  • FIG 1b it is a schematic diagram of a user consuming 3DoF+ immersive media provided by the embodiment of this application.
  • 3DoF+ means that when the virtual scene provided by the immersive media has certain depth information, the user's head can move in a limited space based on 3DoF to watch the screen provided by the media content.
  • FIG. 1c it is a schematic diagram of a user consuming 6DoF immersive video provided by the embodiment of the present application.
  • 6DoF is divided into window 6DoF, omnidirectional 6DoF and 6DoF, where window 6DoF means that the user's rotation and movement on the X-axis and Y-axis are affected. limited, and limited translation in the Z axis; for example, the user cannot see outside the window frame, and the user cannot walk through the window.
  • Omni-directional 6DoF means that the user's rotation and movement on the X-axis, Y-axis and Z-axis are limited. For example, the user cannot freely pass through the three-dimensional 360-degree VR content in the restricted movement area.
  • 6DoF means that users can freely translate along the X-axis, Y-axis, and Z-axis. For example, users can move freely in three-dimensional 360-degree VR content.
  • 6DoF immersive video not only allows users to rotate and consume media content along the X-axis, Y-axis and Z-axis, but also freely move along the X-axis, Y-axis and Z-axis to consume media content.
  • Immersive media content includes video content represented in a three-dimensional (3-Dimension, 3D) space in various forms, for example, three-dimensional video content represented in a spherical form.
  • immersive media content can be VR (Virtual Reality, virtual reality) video content, multi-view video content, panoramic video content, spherical video content or 360-degree video content; therefore, immersive media can also be called VR video, free viewing angle video, panoramic video, spherical video, or 360-degree video.
  • immersive media content also includes audio content synchronized with video content represented in three-dimensional space.
  • FIG. 2a it is a structural diagram of an immersive media system provided by an embodiment of the present application.
  • the immersive media system shown in Figure 2a includes a content production device and a content consumption device
  • the content production device may refer to a computer device used by a provider of immersive media (such as a content producer of immersive content)
  • the computer device may be a terminal , such as smartphones, tablet computers, laptops, desktop computers, smart speakers, smart watches, smart cars, etc.
  • the computer device can also be a server, such as an independent physical server, or a server cluster composed of multiple physical servers or Distributed systems can also be the basis for providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms Cloud server for cloud computing service.
  • a content consumption device may refer to a computer device used by a user (such as a user) who is immersed in media, and the computer device may be a terminal, such as a personal computer, an intelligent mobile device such as a smart phone, or a VR device (such as a VR helmet, VR glasses, etc.).
  • the data processing process of immersive media includes the data processing process on the content production device side and the data processing process on the content consumption device side.
  • the data processing process on the content production device side mainly includes: (1) the acquisition and production process of media content for immersive media; (2) the process of encoding and packaging immersive media.
  • the data processing process on the content consumption side mainly includes: (3) the decapsulation and decoding process of immersive media; (4) the rendering process of immersive media.
  • the transmission process involving immersive media between content production equipment and content consumption equipment can be carried out based on various transmission protocols.
  • the transmission protocols here can include but are not limited to: DASH (Dynamic Adaptive Streaming over HTTP, dynamic Adaptive streaming media transmission) protocol, HLS (HTTP Live Streaming, dynamic code rate adaptive transmission) protocol, SMTP (Smart Media Transport Protocol, intelligent media transmission protocol), TCP (Transmission Control Protocol, transmission control protocol), etc.
  • FIG. 2b it is a schematic diagram of an immersive media transmission solution provided by an embodiment of the present application.
  • the original immersive media is usually chosen to be divided into multiple video segments in space. , respectively encoded and encapsulated, and then transmitted to the client for consumption.
  • the data processing process of the immersive media will be described in detail below in combination with FIG. 2b. First, the data processing process on the content production device side is introduced:
  • the media content of immersive media is obtained by capturing real-world audio-visual scenes through capture devices.
  • the capture device may refer to a hardware component provided in the content production device, for example, the capture device refers to a microphone, a camera, and a sensor of a terminal.
  • the capture device may also be a hardware device that is independent but connected to the content production device, such as a camera connected to a server.
  • the capture device may include but not limited to: audio device, camera device and sensor device.
  • the audio device may include an audio sensor, a microphone, and the like.
  • the camera device may include a common camera, a stereo camera, a light field camera, and the like.
  • Sensing devices may include laser devices, radar devices, and the like.
  • the data of the capture device can be multiple, and these capture devices can be deployed at some specific perspectives in the real space to simultaneously capture audio content and video content from different perspectives in the space, and the captured audio content and video content are separated in time and space are kept in sync.
  • the media content of 3DoF immersive content is recorded by a group of cameras or a camera device with multiple cameras and sensors, and the media content of 6DoF immersive media is mainly in the form of point clouds and light fields captured by camera arrays. content produced.
  • S2 The production process of media content for immersive media.
  • the captured audio content itself is content suitable for audio encoding of immersive media, so no other processing needs to be performed on the captured audio content.
  • the captured video content needs to go through a series of production processes before it can be called content suitable for video encoding of immersive media.
  • the production process can specifically include:
  • splicing refers to splicing the video content captured from these various angles of view into a complete video that can reflect the 360-degree visual panorama of the real space , that is, the stitched video is a panoramic video represented in a three-dimensional space.
  • projection refers to the process of mapping a three-dimensional video formed by splicing onto a two-dimensional (2-Dimension, 2D) image.
  • the 2D image formed by projection is called a projection image; projection methods may include but are not limited to: latitude and longitude Graph projection, regular hexahedron projection.
  • the capture device can only capture panoramic video, after such video is processed by the content production device and transmitted to the content consumption device for corresponding data processing, the user on the content consumption device side can only perform some specific actions (such as head rotation) to watch 360-degree video information, but performing non-specific actions (such as moving the head) cannot obtain corresponding video changes, and the VR experience is not good, so it is necessary to provide additional depth information that matches the panoramic video.
  • some specific actions such as head rotation
  • non-specific actions such as moving the head
  • this involves a variety of production technologies common production technologies include 6DoF production technology, 3DoF production technology and 3DoF+ production technology.
  • Immersive media obtained by using 6DoF production technology and 3DoF+ production technology can include free-view video.
  • free-view video is an immersive media video that is captured by multiple cameras, contains different perspectives, and supports user 3DoF+ or 6DoF interaction.
  • 3DoF+ immersive media is recorded by a set of cameras or a camera with multiple cameras and sensors, and the cameras can usually acquire content in all directions around the center of the device.
  • 6DoF immersive media is mainly made of content in the form of point clouds and light fields captured by camera arrays.
  • the projected image can be encoded directly, or the projected image can be encoded after area encapsulation.
  • Fig. 3a it is a basic block diagram of video coding provided by the embodiment of the present application.
  • Modern mainstream video coding technology taking international video coding standard HEVC (High Efficiency Video Coding), international video coding standard VVC (Versatile Video Coding), and Chinese national video coding standard AVS (Audio Video Coding Standard) as examples, adopts hybrid coding
  • the framework performs the following series of operations and processing on the input original video signal:
  • Block partition structure According to the size of the processing unit, the input image is divided into several non-overlapping processing units, and a similar compression operation is performed on each processing unit. This processing unit is called Coding Tree Unit (CTU), or Largest Coding Unit (LCU). CTU can continue to be more finely divided to obtain one or more basic coding units, called coding units (Coding Unit, CU). Each CU is the most basic element in a coding scheme.
  • Fig. 3b it is a schematic diagram of an input image division adopted by the embodiment of the present application. The following describes various encoding methods that may be used for each CU.
  • Predictive Coding Including Intra(picture) Prediction and Inter(picture) Prediction.
  • a residual video signal is obtained.
  • the content production device needs to select the most suitable one among many possible predictive coding modes for the current CU, and inform the content consumption device.
  • the signal predicted by intra-frame prediction comes from the encoded and reconstructed area in the same image
  • the signal predicted by inter-frame prediction comes from another image that has been encoded and is different from the current image (called a reference image ).
  • Transform&Quantization The residual video signal is transformed into the transform domain through discrete Fourier transform (Discrete Fourier Transform, DFT), discrete cosine transform (Discrete Cosine Transform, DCT) and other transformation operations, are called transformation coefficients.
  • DFT discrete Fourier Transform
  • DCT discrete cosine transform
  • the signal in the transform domain is further subjected to a lossy quantization operation to lose certain information, so that the quantized signal is conducive to compressed expression.
  • the content production device also needs to select one of the transformation methods for the current coding CU, and notify the content playback device.
  • the fineness of quantization is usually determined by the quantization parameter (Quantization Parameter, QP).
  • a larger value of QP means that coefficients with a larger range of values will be quantized to the same output, which usually results in greater distortion, and Lower code rate; on the contrary, the QP value is smaller, which means that the coefficients with a smaller range of values will be quantized to the same output, so it usually brings smaller distortion and corresponds to a higher code rate.
  • Entropy coding or statistical coding: the quantized transform domain signal will be statistically compressed and coded according to the frequency of occurrence of each value, and finally a binary (0 or 1) compressed code stream will be output. At the same time, encoding generates other information, such as selected modes, motion vectors, etc., which also require entropy encoding to reduce the bit rate.
  • Statistical coding is a lossless coding method that can effectively reduce the bit rate required to express the same signal. Common statistical coding methods include variable length coding (VLC, Variable Length Coding) or context-based binary arithmetic coding (CABAC, Content Adaptive Binary Arithmetic Coding).
  • Loop Filtering After the coded image is inversely quantized, inversely transformed, and predictively compensated (reverse operation of 2 to 4 above), the reconstructed decoded image can be obtained. Compared with the original image, the reconstructed image has some information different from the original image due to the influence of quantization, resulting in distortion (Distortion). Filtering operations on the reconstructed image, such as deblocking, sample adaptive offset (Sample Adaptive Offset, SAO) filter or adaptive loop filter (Adaptive Loop Filter, ALF), etc., can effectively reduce Quantizes the amount of distortion produced. Since these filtered reconstructed images will be used as references for subsequent encoded images for predicting future signals, the above filtering operation is also called loop filtering, and filtering operations in the encoding loop.
  • SAO sample adaptive offset
  • ALF adaptive Loop Filter
  • 6DoF Six Degrees of Freedom, six degrees of freedom
  • 6DoF a specific Encoding method (such as point cloud encoding) for encoding.
  • the file can be a media file of immersive media formed by a media file or a media fragment, and use media presentation description (MPD) to record the metadata of the media file resource of the immersive media according to the requirements of the file format of the immersive media, here Metadata is a general term for information related to the presentation of immersive media.
  • the metadata may include description information about media content, description information about windows, and signaling information related to the presentation of media content. As shown in FIG. 2a, the content production device stores media presentation description information and media file resources formed after the data processing process.
  • the content consumption device can obtain immersive media media file resources and corresponding media presentation description information dynamically from the content production device through the recommendation of the content production device or adaptively according to the needs of the user on the content consumption device side.
  • the eye/body tracking information determines the user's orientation and position, and then dynamically requests the content production device to obtain corresponding media file resources based on the determined orientation and position.
  • Media file resources and media presentation description information are transmitted from the content production device to the content consumption device through a transmission mechanism (such as DASH, SMT).
  • the decapsulation process of the content consumption device is opposite to the encapsulation process of the content production device.
  • the content consumption device decapsulates the acquired media file resources according to the file format requirements of immersive media, and obtains the audio code stream and video code stream.
  • the decoding process of the content consumption device is opposite to the encoding process of the content production device.
  • the content consumption device decodes the audio code stream to restore the audio content
  • the content consumption device decodes the video code stream to obtain the video content.
  • the decoding process of the video code stream by the content consumption device may include the following: 1 Decoding the video code stream to obtain a planar projection image. 2The projection image is reconstructed according to the media presentation description information to convert it into a 3D image.
  • the reconstruction process here refers to the process of reprojecting the two-dimensional projection image into a 3D space.
  • the content consumption device first performs entropy decoding to obtain various mode information and quantized transformation coefficients. Each coefficient is repeatedly quantized and transformed to obtain a residual signal.
  • the predicted signal corresponding to the CU can be obtained, and after the two are added together, the reconstructed signal can be obtained.
  • the reconstructed value of the decoded image needs to undergo a loop filtering operation to generate the final output signal.
  • the content consumption device renders the audio content obtained by audio decoding and the 3D image obtained by video decoding according to the metadata related to rendering and window in the media subsidence description information. After the rendering is completed, the playback and output of the 3D image is realized.
  • the content consumption device mainly renders the 3D image based on the current viewpoint, disparity, depth information, etc. The image is rendered.
  • the viewpoint refers to the viewing position of the user
  • the parallax refers to the visual difference caused by the user's binoculars or due to movement
  • the window refers to the viewing area.
  • a data box refers to a data block or object including metadata, that is, a data box includes metadata of corresponding media content. It can be seen from the above data processing process of immersive media that after encoding the immersive media, the encoded immersive media needs to be encapsulated and transmitted to the user. Immersive media in the embodiment of this application mainly refers to free-view video.
  • Atlas information can be obtained only with camera parameters, and the positions of texture images and depth images in plane frames are also
  • the large-scale atlas information data box can be used to indicate the relevant parameter information, thereby omitting the rest of the atlas information in the atlas track.
  • camera_count indicates the number of all cameras that capture immersive media
  • padding_size_depth indicates the width of the guard band used when encoding the depth image
  • padding_size_texture indicates the guard band used when encoding the texture image Width
  • camera_id indicates the camera identifier of a camera in a viewing angle
  • camera_resolution_x indicates the resolution width of a texture image and a depth image collected by a camera
  • camera_resolution_y indicates the resolution height of a texture image and a depth image collected by a camera
  • depth_downsample_factor indicates the downsampling of a depth image Sampling multiple factor, the actual resolution width and height of the depth image is 1/2 depth_downsample_factor of the camera acquisition resolution width and height
  • depth_vetex_x indicates the horizontal axis of the offset of the upper left vertex of the depth image relative to the origin of the plane frame (the upper left vertex of the plane frame) Component
  • depth_vetex_y indicates the
  • the large-scale atlas information data box indicates the layout information of the texture image and depth image in the free-view video frame, and gives the relevant camera parameters For example, camera_resolution_y and camera_resolution_x, etc., but the above only considers the scene where the free-view video is encapsulated in a single track, and does not consider the scene where the free-view video is packaged in multiple tracks.
  • the above-mentioned large-scale atlas information data box indicates the arrangement information of the texture map and depth map in the free-view video and related camera parameters, but only considers the case of encapsulating the free-view video into a single track, and does not consider In the case of multi-track packaging, the camera parameters indicated in the large-scale atlas information data box cannot be used as the basis for content consumption devices to select images from different perspectives for decoding and consumption. That is to say, according to the above-mentioned large-scale atlas information data box camera parameters, the content consumption device cannot know which image is suitable for the current user location information, which brings inconvenience to the decoding of the content consumption device.
  • an embodiment of the present application provides a data processing scheme for immersive media, in which the immersive media is encapsulated into M tracks, and the immersive media is composed of images taken by N cameras at different viewing angles Yes, these M tracks belong to the same track group, which realizes a scene where immersive video is encapsulated into multiple tracks; in addition, the content production device generates a free-view information data box for each track, and corresponds to the i-th track
  • the free viewing angle information data box in indicates the viewing angle information corresponding to the i-th track, such as the specific viewing angle position of the camera, and then when the content consumption device decodes and displays the image in the track according to the viewing angle information corresponding to the i-th track, it can ensure that The decoded and displayed video matches the user's location more closely, improving the rendering effect of immersive video.
  • FIG. 4 it is a schematic flowchart of a data processing method for immersive media provided by an embodiment of the present application.
  • the data processing method described in FIG. 4 may be executed by a content consumption device, specifically, may be executed by a processor of the content consumption device.
  • the data processing method shown in Figure 4 may include the following steps:
  • the immersive media is composed of images taken by N cameras at different viewing angles, and the immersive media is encapsulated into M tracks, and one track encapsulates images from at least one camera, and these M tracks belong to the same track group .
  • the i-th track refers to a selected one of the M tracks, and how to select the i-th track from the M tracks will be described in detail later.
  • the M tracks encapsulating the immersive video are associated through the information of the free-viewing track group encapsulated in each track.
  • the i-th track is encapsulated with the free-view track group information corresponding to the i-th track.
  • the free-view track group information is used to indicate that the i-th track and other tracks that encapsulate immersive video belong to the same track group.
  • the free-view track group information encapsulated in the i-th track can be obtained by expanding the track group data box, and the syntax of the free-view track group information packaged in the i-th track can be expressed as the following code segment 2:
  • the free view track group information is obtained by extending the track group data box, and is identified by the "a3fg" track group type. Among all tracks containing TrackGroupTypeBox of type "afvg", tracks with the same group ID belong to the same track group.
  • the i-th track encapsulates the images collected by the k cameras of the aforementioned N cameras, and the images collected by any camera can be at least one of texture images or depth images.
  • One camera corresponds to one identification information, and each identification information is stored in the first camera identification field camera_id of the free-view track group information. It should be understood that one camera corresponds to one identification information, and the i-th track includes k cameras. Therefore,
  • the free track group information includes k first camera identification fields camera_id.
  • the free track group information also includes the image type field depth_texture_type, which is used to indicate the image type of the image captured by the jth camera, and j is greater than 0 and less than k, that is to say, an image type field depth_texture_type is used to indicate a camera
  • the free track group information includes k image type fields.
  • the image type field depth_texture_type indicating the image type of the image captured by the jth camera can be specifically referred to in Table 1 below:
  • the image type field depth_texture_type when the image type field depth_texture_type is 1, it indicates that the image type of the image captured by the jth camera is a texture image; when the image type field depth_texture_type is 2, it indicates that the image type of the image captured by the jth camera is depth Image; when the image type field depth_texture_type is 3, it indicates that the image type of the image captured by the jth camera is texture image and depth image.
  • the free view angle information data box AvsFreeViewInfoBox corresponding to the i-th track includes the view angle information corresponding to the i-th track
  • code segment 3 which is the syntax representation of the free view angle information data box corresponding to the i-th track
  • code segment 3 is as follows:
  • the viewing angle information corresponding to the i-th track may include a video stitching layout indication field stitching_layout, which is mainly used to indicate whether the texture image and depth image included in the i-th track are stitched and coded; specifically, when the video stitching layout
  • the indication field stitching_layout is the first data value, it indicates that the texture image and depth image included in the i-th track are stitched and coded
  • the video stitching layout indication field stitching_layout is the second value, it indicates the texture included in the i-th track Images and depth images are coded separately. For example, assuming that immersive media refers to 6DoF video, the first value is 0, and the second value is 1, then the video stitching layout indication field stitching_layout in the i-th track can be shown in Table 2 below:
  • the angle of view information corresponding to the i-th track also includes the camera model field camera_model, which is used to indicate the camera model of k cameras in the i-th track.
  • k represents the depth image and/or texture in the i-th track The number of cameras from which the image is sourced.
  • the angle of view information corresponding to the i-th track can also include the camera number field camera_count, which is used to store the number of cameras from which the depth image and/or texture image in the i-th track is assumed to indicate for k.
  • the camera model field camera_model when the camera model field camera_model is the third value, it indicates that the camera model of the jth camera is the first model; when the camera model field camera_model is the fourth value, it indicates that the camera model of the jth camera is the second model .
  • the first model can refer to the pointer hole model
  • the second model can refer to the fisheye model.
  • the third value is 0, the fourth value is 1, the immersive media is 6DoF video, and the camera model in the viewing angle information corresponding to the i-th track
  • the field camera_model can be seen in Table 3 below:
  • the angle of view information corresponding to the i-th track also includes the guard band width field texture_padding_size of the texture image, and the guard band width field depth_padding_size of the depth image.
  • the guard band width field of the texture image is used to store the guard band width used when encoding the texture image in the i-th track
  • the guard band width of the depth image is used to store the guard band width used when encoding the depth image in the i-th track .
  • the angle of view information corresponding to the i-th track also includes the second camera identification field camera_id, which is used to store the identification information of the j-th camera in the i-th track.
  • the texture image and depth image source included are k cameras, and the value of j is greater than or equal to 0 and less than k. That is to say, a second camera identification field camera_id stores the identification information of any one of the k cameras, therefore, k camera identification fields are required to store the identification information of the k cameras.
  • the second camera identification field camera_id here has the same function as the first camera identification field camera_id in the aforementioned free-view track group information, both of which are used to store the identification information of the j-th camera in the i-th track .
  • the angle of view information corresponding to the i-th track also includes a camera attribute information field, which is used to store the camera attribute information of the j-th camera, and the camera attribute information of the j-th camera can include the horizontal axis of the j-th camera position The value of the component, the value of the vertical axis component and the value of the vertical axis component, the value of the horizontal axis component and the value of the vertical axis component of the focal length of the jth camera, and the resolution width and height of the image captured by the jth camera.
  • the camera attribute information field may specifically include: 1) the horizontal axis component field camera_pos_x of the camera position, which is used to store the horizontal axis component value (also called the x component value) of the jth camera position; 2) the camera position
  • the vertical axis component field camera_pos_y is used to store the value of the vertical axis component of the jth camera position (also called the value of the y component); 3) the vertical axis component field of the camera position camera_pos_z is used to store the value of the jth camera position
  • the value of the vertical axis component also called the value of the z component
  • the field focal_length_x of the horizontal axis component of the focal length of the camera which is used to store the value of the horizontal axis component of the jth camera focal length (also called the value of the x component); 5)
  • the field focal_length_y of the vertical axis component of the focal length of the camera is used to store the value of the vertical axis component of the j
  • one camera attribute information field is used to store the camera attribute information of a camera, therefore, in the viewing angle information corresponding to the i-th track K camera attribute information fields are included for storing the camera attribute information of k cameras.
  • one camera attribute field includes the above 1)-7), and k pieces of camera attribute information include k pieces of the above 1)-7).
  • the angle of view information corresponding to the i-th track also includes an image information field, which is used to store the image information of the image captured by the j-th camera.
  • the image information may include at least one of the following: a downsampling factor of the depth image, an offset of the upper left vertex of the depth image relative to the origin of the plane frame, or an offset of the upper left vertex of the texture image relative to the origin of the plane frame.
  • the image information field may specifically include: 1) the depth_downsample_factor field for downsampling the depth image, which is used to store the downsampling factor for the depth image; 2) the horizontal axis offset field texture_vetex_x for the top left vertex of the texture image, which is used to store the texture The horizontal axis component of the offset of the upper left vertex of the image relative to the origin of the plane frame; 3) the vertical axis offset field of the upper left vertex of the texture image texture_vetex_y, which is used to store the offset vertical axis component of the upper left vertex of the texture image relative to the origin of the plane frame; 4 ) The horizontal axis offset field of the upper left vertex of the depth image depth_vetex_x is used to store the horizontal axis offset of the upper left vertex of the depth image relative to the origin of the plane frame; 5) The vertical axis offset field of the upper left vertex of the depth image depth_vetex_y is used to store the upper left of the depth image
  • an image information field is used to store the image information of an image captured by a camera, and the i-th track includes depth images and/or texture images captured by k cameras, so the free viewing angle information corresponding to the i-th track Including k image information fields.
  • the angle of view information corresponding to the i-th track also includes a custom camera parameter field camera_parameter, which is used to store the f-th custom camera parameter of the j-th camera, where f is an integer greater than or equal to 0 and less than h , h represents the number of custom camera parameters of the jth camera, and the number of custom camera parameters of the jth camera can be stored in the field para_num of the number of custom camera parameters of the viewing angle information.
  • the custom camera parameter field corresponding to the jth camera The quantity can be h pieces.
  • the viewing angle information corresponding to the i-th track includes k*h custom camera parameter fields.
  • the angle of view information corresponding to the i-th track also includes a custom camera parameter type field para_type, which is used to store the type of the f-th custom camera parameter of the j-th camera. It should be noted that the type of a custom camera parameter of the jth camera is stored in a custom camera parameter type field. Since the jth camera corresponds to h custom camera parameters, the number of custom camera parameter type fields for h. Similar to 7, the angle of view information corresponding to the i-th track includes k*h custom camera parameter type fields.
  • the angle of view information corresponding to the i-th track also includes a custom camera parameter length field para_length, which is used to store the length of the f-th custom camera parameter of the j-th camera. Since the length of a custom camera parameter of the jth camera is stored in a custom camera parameter length field, and because the jth camera includes h custom camera parameters, the number of custom camera parameter length fields is h.
  • obtaining the free view angle information data box corresponding to the i-th track in S401 is realized based on the signaling description file sent by the content production device.
  • the signaling description file corresponding to the immersive media is acquired, the signaling description file includes the free-view camera descriptor corresponding to the immersive media, and the free-view camera descriptor is used to record the Camera attribute information, video clips in any track are composed of texture images and/or depth images in any track; the free-view camera descriptor is encapsulated in the adaptive media presentation description file of the media data set level, or the signaling description file is encapsulated in the presentation level of the media presentation description file.
  • the free view camera descriptor can be expressed as AvsFreeViewCamInfo, which is a SupplementalProperty element, and its @schemeIdUri attribute is "urn:avs:ims:2018:av3l".
  • AvsFreeViewCamInfo which is a SupplementalProperty element
  • @schemeIdUri attribute is "urn:avs:ims:2018:av3l”.
  • each element and attribute in the free-view camera descriptor can be shown in Table 4 below:
  • the content consumption device acquires the free view angle information data box corresponding to the i-th track based on the signaling description file.
  • obtaining the free view information data box corresponding to the i-th track includes: based on the camera attribute information corresponding to the image in each track recorded in the free view camera descriptor, select from the N cameras that correspond to the user's location Candidate cameras whose location information matches; send to the content production device a first resource request for acquiring images taken by the candidate camera, the first resource request is used to instruct the content production device to use the free view angle information data box corresponding to each track in the M tracks Angle of view information, select the i-th track from M tracks and return the free angle of view information data box corresponding to the i-th track, the i-th track encapsulates the image taken by the candidate camera; the i-th track returned by the receiving content production device corresponds to The free view information data box. In this way, the content consumption device only needs to obtain the free-view information data boxes corresponding to the required
  • each track can include a texture image and a depth image of a perspective, that is to say, a track encapsulates a texture from a camera Image and depth image, texture image and depth image in a track form a video clip, so it can be understood that a track encapsulates a video clip from a camera at a viewing angle.
  • the free-view video is encapsulated into three tracks, and the texture image and depth image of the video in each track form a video clip, so the free-view video includes three video clips, which are respectively represented as Representation1, Representation2 and Representation3;
  • the content production device In the signaling generation link, the content production device generates the signaling description file corresponding to the free-view video according to the view information in the free-view information data box corresponding to each track, and the free-view camera description can be carried in the signaling description file son.
  • the free-view video free-view camera descriptor records the camera attribute information corresponding to 3 video clips as follows:
  • the content production device sends the signaling description file to the content consumption device
  • the content consumption device selects the video clips from Camera2 and Camera3 and requests the content production device according to the signaling description file and user bandwidth, according to the location information of the user and the camera attribute information in the signaling description file;
  • the content production device sends the free view information data boxes of the tracks of the video clips from Camera2 and Camera3 to the content consumption device, and the content consumption device initializes the decoder according to the acquired free view information data boxes corresponding to the tracks, and decodes the corresponding Video clips and consumption.
  • obtaining the free view information data box corresponding to the i-th track includes: sending a second resource request to the content production device based on the signaling description file and user bandwidth, and the second resource request is used to instruct the content production device to return M free perspective information data boxes of M tracks, one track corresponds to one free perspective information data box; according to the perspective information in the free perspective information data box corresponding to each track and the user's location information, from the M free perspective information data Get the free viewing angle information data box corresponding to the i-th track in the box.
  • the content consumption device obtains the free view information data boxes of all tracks, it does not decode and consume images in all tracks, but only decodes the image in the i-th track that matches the user's current location, saving decoding resources . For example:
  • each track can include a texture image and a depth image of a view, that is to say, each track encapsulates a camera from a view
  • each track encapsulates a camera from a view
  • the free-view video is encapsulated into three tracks of video clips respectively represented as Representation1, Representation2, and Representation3;
  • the content production device generates the signaling description file of the free-view video according to the viewing angle information corresponding to each track, as follows:
  • the content production device sends the signaling description file in (2) to the content consumption device;
  • the content consumption device requests the free-view information data boxes corresponding to all tracks from the content production device according to the signaling description file and user bandwidth; assuming that the free-view video is encapsulated into three tracks, Track1, Track2, and Track3 each
  • the camera attribute information corresponding to the video clip encapsulated in each track can be as follows:
  • the content consumption device selects the video clips encapsulated in the tracks of track2 and track3 for decoding and consumption according to the obtained viewing angle information in the free viewing angle information data box corresponding to all tracks and the position information currently viewed by the user.
  • S402. Decode the image encapsulated in the i-th track according to the viewing angle information in the free viewing angle information data box.
  • decoding and displaying the image encapsulated in the i-th track according to the angle of view information in the free angle of view information data box may include: initializing the decoder according to the angle of view information corresponding to the i-th track; The relevant instruction information decodes the image encapsulated in the i-th track.
  • an immersive medium is packaged into M tracks.
  • the immersive medium is composed of images taken by N cameras at different viewing angles.
  • One track may include images from at least one camera.
  • the M tracks are Belonging to the same track group, a scene in which an immersive video is encapsulated into multiple tracks is realized; in addition, the content production device generates a free-view information data box for each track, which is indicated by the free-view information data box corresponding to the i-th track.
  • the viewing angle information corresponding to the i-th track such as the specific viewing angle position of the camera, and then when the content consumption device decodes and displays the image in the track according to the viewing angle information corresponding to the i-th track, it can ensure that the decoded and displayed video is consistent with the user's location. The position is more matched, improving the presentation effect of immersive video.
  • FIG. 5 it is a schematic flowchart of a data processing method for immersive media provided by an embodiment of the present application.
  • the data processing method for immersive media shown in FIG. 5 can be executed by a content production device, specifically, by a processor of the content production device.
  • the data processing method of the immersive media shown in Figure 5 may include the following steps:
  • the immersive media is composed of images taken by cameras at different viewing angles, one track encapsulates images from at least one camera, and the M tracks belong to the same track group .
  • w is an integer greater than or equal to 1
  • w images are encapsulated in the i-th track, and these w images can be at least one of depth images and texture images taken by a camera at one or several viewing angles kind.
  • the code segment 3 in the embodiment of FIG. 4 it is specifically introduced to generate the free view angle information data box corresponding to the i-th track according to the encapsulation process of w images in the i-th track, which may specifically include:
  • the free-view information data box corresponding to the i-th track includes free-view track group information, which can indicate that the i-th track belongs to the same track group as other tracks encapsulating immersive media.
  • the free viewing angle track group information includes the first camera identification field camera_id and the image type field depth_texture_type.
  • the free viewing angle information data box corresponding to the ith track is generated according to the packaging process of the w images in the ith track, including: Determine that w images come from k cameras, k is an integer greater than 1; the identification information of the j camera in the k cameras is stored in the first camera identification field camera_id, j is greater than 0 and less than k; according to the The image type of the image captured by the jth camera determines the image type field depth_texture_type, and the image type includes any one or more of a depth image and a texture image.
  • the identification information of a camera is stored in a first camera identification field camera_id, if the images contained in the i-th track come from k cameras, then the free-view track group information may include k first camera identification fields camera_id. Similarly, the free view track group includes k image type fields depth_texture_type.
  • determining the image type field according to the image type of the image collected by the jth camera includes: when the image type of the image collected by the jth camera is a texture image, then setting the image type field to a second value; The image type of the image collected by the first camera is a depth image, then the image type field is set to the third value; when the image type of the image collected by the jth camera is a texture image and a depth image, the image type field is set to the fourth value value.
  • the second value may be 1
  • the third value may be 2
  • the fourth value may be 3
  • the image type field may refer to Table 2 in the embodiment of FIG. 4 .
  • the viewing angle information corresponding to the i-th track includes the video splicing layout indication field stitching_layout.
  • the free viewing angle information data box corresponding to the i-th track is generated according to the encapsulation process of the w images in the i-th track, including: if the i-th If the texture image and depth image contained in the track are coded by splicing, set the video layout indication field stitching_layout to the first value; if the texture image and depth image contained in the i-th track are coded separately, set the video layout indication
  • the field stitching_layout is set to the second value. Wherein, the first value may be 0, the second value may be 1, and the video mosaic layout indication field may refer to Table 3 in the embodiment of FIG. 4 .
  • the angle of view information corresponding to the i-th track also includes the camera model field camera_model; in S502, the free angle of view information data box corresponding to the i-th track is generated according to the encapsulation process of the w images in the i-th track, including: when w images are collected When the camera model of the camera belonging to is the first model, set the camera model field camera_model to the first value; when the camera model of the camera that captures w images belongs to the second model, set the camera model field camera_model to the second value .
  • the first value may be 0, the second value may be 1, and the camera model field may refer to Table 4 in the embodiment of FIG. 4 .
  • the angle of view information corresponding to the i-th track also includes the protection bandwidth field texture_padding_size of the texture image and the protection bandwidth field depth_padding_size of the depth image.
  • a free-view information data box corresponding to the i-th track is generated, including: obtaining the guard band width of the texture image used when encoding the texture image in the i-th track, and Store the guard band width of the texture image in the guard band width field texture_padding_size of the texture image; obtain the guard band width of the depth image used when encoding the depth image in the i-th track, and store the guard band width of the depth image in depth The guard band width field depth_padding_size of the image.
  • the viewing angle information corresponding to the i-th track also includes a second camera identification field camera_id and a camera attribute information field.
  • a free viewing angle information data box corresponding to the i-th track is generated according to the encapsulation process of w images in the i-th track, Including: storing the identification information of the jth camera in the second camera identification field camera_id, where j is greater than or equal to 0 and less than k, and k represents the number of w image source cameras; obtaining the camera attribute information of the jth camera, and The obtained camera attribute information is stored in the camera attribute information field.
  • the camera attribute information of the jth camera includes any one or more of the following: the value of the horizontal axis component, the value of the vertical axis component, and the value of the vertical axis component of the position of the jth camera, the focal length of the jth camera The value of the horizontal axis component and vertical axis component of , and the resolution width and height of the image captured by the jth camera.
  • the camera attribute information field may specifically include: 1) the horizontal axis component field camera_pos_x of the camera position, which is used to store the horizontal axis component value (also called the x component value) of the jth camera position; 2) the camera position
  • the vertical axis component field camera_pos_y of the camera position is used to store the vertical axis component value of the jth camera position (also called the y component value); 3) the vertical axis component field of the camera position camera_pos_z is used to store the jth camera position
  • the value of the vertical axis component also called the value of the z component
  • the field focal_length_x of the horizontal axis component of the focal length of the camera which is used to store the value of the horizontal axis component of the jth camera focal length (also called the value of the x component) ;5)
  • the field focal_length_y of the vertical axis component of the focal length of the camera is used to store the value of the vertical axis component of the jth camera
  • one camera attribute information field is used to store the camera attribute information of a camera, therefore, in the viewing angle information corresponding to the i-th track K camera attribute information fields are included for storing camera attribute information of k cameras.
  • one camera attribute field includes the above 1)-7), and k pieces of camera attribute information include k pieces of the above 1)-7).
  • the angle of view information corresponding to the i-th track also includes an image information field.
  • the free angle of view information data box corresponding to the i-th track is generated according to the encapsulation process of the w images, including: acquiring the image of the image captured by the j-th camera information, and store the obtained image information in the image information field; wherein, the image information includes one or more of the following: the multiplication factor of the downsampling of the depth image, the offset of the upper left vertex of the depth image relative to the origin of the plane frame, The offset of the top left vertex of the texture image relative to the origin of the plane frame.
  • the image information field may specifically include: 1) the depth_downsample_factor field for downsampling the depth image, which is used to store the downsampling factor for the depth image; 2) the horizontal axis offset field texture_vetex_x for the top left vertex of the texture image, which is used to store the texture The horizontal axis component of the offset of the upper left vertex of the image relative to the origin of the plane frame; 3) the vertical axis offset field of the upper left vertex of the texture image texture_vetex_y, which is used to store the offset vertical axis component of the upper left vertex of the texture image relative to the origin of the plane frame; 4 ) The horizontal axis offset field of the upper left vertex of the depth image depth_vetex_x is used to store the horizontal axis offset of the upper left vertex of the depth image relative to the origin of the plane frame; 5) The vertical axis offset field of the upper left vertex of the depth image depth_vetex_y is used to store the upper left of the depth image
  • an image information field is used to store the image information of an image captured by a camera, and the i-th track includes depth images and/or texture images captured by k cameras, so the free viewing angle information corresponding to the i-th track Including k image information fields.
  • the angle of view information corresponding to the i-th track also includes a custom camera parameter field, a custom camera parameter type field, and a custom camera parameter length field; Free viewing angle information data box, including: obtaining the fth custom camera parameter, and storing the fth custom camera parameter in the custom camera parameter field, f is an integer greater than or equal to 0 and less than h, and h represents the jth The number of custom camera parameters of the camera; determine the parameter type of the fth custom camera parameter and the length of the fth custom camera parameter, and store the parameter type of the fth custom camera parameter in the custom camera parameter type field, and store the length of the fth custom camera parameter in the custom camera parameter length field.
  • the custom camera parameter field corresponding to the jth camera The quantity can be h pieces. And because each camera corresponds to h custom camera parameters, and the i-th track includes k cameras, therefore, the viewing angle information corresponding to the i-th track includes k*h custom camera parameter fields. It should be noted that since the number of custom camera parameters of the jth camera is h, one custom camera parameter field is used to store a custom camera parameter, therefore, the custom camera parameter field corresponding to the jth camera The quantity can be h pieces.
  • the viewing angle information corresponding to the i-th track includes k*h custom camera parameter fields. Since the length of a custom camera parameter of the jth camera is stored in a custom camera parameter length field, and because the jth camera includes h custom camera parameters, the number of custom camera parameter length fields is h.
  • the content production device can also generate a signaling description file according to the viewing angle information in the free viewing angle information data box corresponding to the M tracks.
  • the signaling description file includes the free viewing angle camera descriptor corresponding to the immersive media.
  • the free viewing angle camera descriptor is used for Record the camera attribute information corresponding to the image in each track; the free-view camera descriptor is encapsulated in the adaptive set level of the immersive media media presentation description file, or the signaling description file is encapsulated in the presentation level of the media presentation description file .
  • the content production device may send the signaling description file to the content consumption device, so that the content consumption device obtains the free view information data box corresponding to the i-th track according to the signaling description file.
  • the content production device sends the signaling description file to the content consumption device to instruct the content consumption device to, based on the camera attribute information corresponding to the image in each track recorded in the free view camera descriptor, start from Select a candidate camera that matches the user's location from among the cameras, and send a first resource request to obtain an image from the candidate camera; in response to the first resource request, according to the free view angle information data box corresponding to each track in the M tracks
  • the viewing angle information in selects the i-th track from the M tracks and sends the free viewing angle information data box corresponding to the i-th track to the content consumption device, and the i-th track includes images from candidate cameras.
  • the content production device sends the signaling description file to the content consumption device to instruct the content consumption device to send a second resource request according to the signaling description file and user bandwidth; in response to the second resource request, Send the M free viewing angle information data boxes corresponding to M tracks to the content consumption device to instruct the content consumption device to select from the M free viewing angles according to the viewing angle information in the free viewing angle information data box corresponding to each track and the user's location information. Obtain the free viewing angle information data box corresponding to the i-th track in the information data box.
  • one immersive media is packaged into M tracks.
  • the immersive video is composed of images collected by cameras at different angles of view.
  • the M tracks belong to the same track group, and one track encapsulates images from at least one
  • the image of the camera realizes a scene where an immersive video is encapsulated into multiple tracks; in addition, the content production device generates a free view angle information data box for each track according to the encapsulation process of the image in each track, and corresponds to the i-th track
  • the free viewing angle information data box in indicates the viewing angle information corresponding to the i-th track, such as the specific viewing angle position of the camera, and then when the content consumption device decodes and displays the image in the track according to the viewing angle information corresponding to the i-th track, it can ensure that The decoded and displayed video matches the user's location more closely, improving the rendering effect of immersive video.
  • the embodiment of the present application provides a data processing device for immersive media.
  • the data processing device for immersive media may be a computer program (including program code) running in a content consumption device. ), for example, the data processing device for immersive media may be an application software in a content consumption device.
  • FIG. 6 it is a schematic structural diagram of a data processing device for immersive media provided by an embodiment of the present application.
  • the data processing device shown in Figure 6 can run the following units:
  • the obtaining unit 601 is configured to obtain a free view angle information data box corresponding to the i-th track of the immersive media, the free view angle information data box includes the view angle information corresponding to the i-th track, i is greater than or equal to 1 and An integer less than or equal to M;
  • the immersive media is composed of images taken by N cameras at different viewing angles, the immersive media is encapsulated into M tracks, and one track encapsulates images from at least one camera, the M orbitals belong to the same orbital group, and both N and M are integers greater than 1;
  • the processing unit 602 is configured to decode the image encapsulated in the i-th track according to the viewing angle information in the free viewing angle information data box.
  • images from k cameras among the N cameras are encapsulated in the i-th track, and k is an integer greater than 0; the i-th track is also encapsulated in the i-th Free perspective track group information corresponding to the track, the free perspective track group information is used to indicate that the ith track belongs to the same track group as other tracks that encapsulate the immersive media; the free perspective track group information includes the first camera identification field and image type field;
  • the first camera identification field is used to store the identification information of the jth camera among the k cameras
  • the image type field is used to indicate the image type to which the image collected by the jth camera belongs, and the image type At least one of a texture image or a depth image is included.
  • the viewing angle information corresponding to the i-th track includes a video mosaic layout indication field; when the video mosaic layout indication field is the first value When , it indicates that the texture image and depth image included in the i-th track are spliced and coded; when the video splicing layout indication field is the second value, it indicates that the texture image and depth image included in the i-th track Images are encoded separately.
  • the angle of view information corresponding to the i-th track further includes a camera model field; when the camera model field is a third value, it indicates that the camera model to which the j cameras belong is the first model; when the When the camera model field is the fourth value, it indicates that the camera model to which the j cameras belong is the second model.
  • the viewing angle information corresponding to the ith track further includes a guard band width field of a texture image and a guard band width field of a depth image, and the guard band width field of the texture image is used to store the ith
  • the guard band width used when encoding the texture image in the track, and the guard band width field of the depth image is used to store the guard band width used when encoding the depth image in the ith track.
  • the viewing angle information corresponding to the i-th track further includes a second camera identification field and a camera attribute information field; the second camera identification field is used to store the identification information of the j-th camera;
  • the camera attribute information field is used to store the camera attribute information of the jth camera, and the camera attribute information of the jth camera includes at least one of the following: the horizontal axis component value of the jth camera position, The value of the vertical axis component and the value of the vertical axis component, the value of the horizontal axis component and the value of the vertical axis component of the focal length of the jth camera, and the resolution width and height of the image captured by the jth camera.
  • the viewing angle information corresponding to the i-th track further includes an image information field
  • the image information field is used to store image information of an image captured by the j-th camera
  • the image information includes any of the following : The multiplication factor of the depth image downsampling, the offset of the upper left vertex of the depth image relative to the origin of the plane frame, or the offset of the upper left vertex of the texture image relative to the origin of the plane frame.
  • the angle of view information corresponding to the ith track further includes a custom camera parameter field, a custom camera parameter type field, and a custom camera parameter length field;
  • the custom camera parameter field is used to store the The fth custom camera parameter of the jth camera, f is an integer greater than or equal to 0 and less than h, and h represents the number of custom camera parameters of the jth camera;
  • the custom camera parameter type field is used for Store the parameter type to which the fth custom camera parameter belongs;
  • the custom camera parameter length field is used to store the length of the fth custom camera parameter.
  • the acquiring unit 601 is also configured to:
  • the signaling description file includes the free-view camera descriptor corresponding to the immersive media, and the free-view camera descriptor is used to record the camera corresponding to the video segment in each track Attribute information, the video segment in the track is composed of texture images and depth images included in the track; the free-view camera descriptor is encapsulated in the adaptive set level of the media presentation description file of the immersive media , or the signaling description file is encapsulated in the presentation level of the media presentation description file.
  • the acquisition unit 601 performs the following operations when acquiring the free perspective information data box corresponding to the i-th track of the immersive media:
  • the first resource request is used to instruct the content production device according to the free viewing angle corresponding to each track in the M tracks Viewing angle information in the information data box, select the i-th track from the M tracks and return the free viewing angle information data box corresponding to the i-th track, the texture image and depth image encapsulated in the i-th track At least one of them is from the candidate camera; receiving a free-view information data box corresponding to the i-th track returned by the content production device.
  • the obtaining unit 601 when the obtaining unit 601 obtains the free view angle information data box corresponding to the ith track of the immersive media, it performs the following operations:
  • the second resource request is used to instruct the content production device to return M free-view information data boxes of the M tracks, one The track corresponds to a free view information data box;
  • the free angle of view information data box corresponding to the i-th track is obtained from the M free angle of view information data boxes.
  • each step involved in the data processing method for immersive media shown in FIG. 4 may be executed by each unit in the data processing apparatus for immersive media shown in FIG. 6 .
  • S401 described in FIG. 4 may be performed by the acquiring unit 601 in the data processing device shown in FIG. 6
  • S402 may be performed by the processing unit 602 in the data processing device shown in FIG. 6 .
  • each unit in the data processing device for immersive media shown in FIG. can be further divided into a plurality of functionally smaller units, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application.
  • the above-mentioned units are divided based on logical functions.
  • the functions of one unit may also be realized by multiple units, or the functions of multiple units may be realized by one unit.
  • the immersive media-based data processing apparatus may also include other units.
  • these functions may also be implemented with the assistance of other units, and may be implemented cooperatively by multiple units.
  • a general-purpose computing device such as a computer including processing elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM) and storage elements.
  • CPU central processing unit
  • RAM random access storage medium
  • ROM read-only storage medium
  • Run a computer program capable of executing the steps involved in the corresponding method as shown in Figure 4, to construct a data processing device for immersive media as shown in Figure 6, and to realize the data of the immersive media in the embodiment of the present application Approach.
  • the computer program can be recorded in, for example, a computer-readable storage medium, loaded into the above-mentioned computing device through the computer-readable storage medium, and run there.
  • an immersive medium is packaged into M tracks.
  • the immersive medium is composed of images taken by N cameras at different viewing angles.
  • One track may include images from at least one camera.
  • the M tracks are Belonging to the same track group, a scene in which an immersive video is encapsulated into multiple tracks is realized; in addition, the content production device generates a free-view information data box for each track, which is indicated by the free-view information data box corresponding to the i-th track.
  • the viewing angle information corresponding to the i-th track such as the specific viewing angle position of the camera, and then when the content consumption device decodes and displays the image in the track according to the viewing angle information corresponding to the i-th track, it can ensure that the decoded and displayed video is consistent with the user's location. The position is more matched, improving the presentation effect of immersive video.
  • this embodiment of the present application provides another data processing device for immersive media.
  • the data processing device for immersive media may be a computer running in a content production device
  • the program (including program code), for example, the data processing device for the immersive media may be an application software in the content production device.
  • FIG. 7 it is a schematic structural diagram of another data processing device for immersive media provided by an embodiment of the present application.
  • the data processing device shown in Figure 7 can run the following units:
  • the encapsulation unit 701 is configured to encapsulate the immersive media into M tracks, the immersive media is composed of images taken by N cameras at different viewing angles, one track encapsulates images from at least one camera, and the M
  • the orbitals belong to the same orbital group, and both N and M are integers greater than or equal to 1;
  • the generating unit 702 is configured to generate a free viewing angle information data box corresponding to the i-th track according to the encapsulation process of the w images in the i-th track, the free viewing angle information data box including the viewing angle information corresponding to the i-th track ; 1 ⁇ i ⁇ M, w ⁇ 1.
  • the free-view information data box corresponding to the i-th track includes free-view track group information
  • the free-view track group information includes a first camera identification field and an image type field
  • the encapsulation unit 701 When generating the free view angle information data box corresponding to the i-th track according to the encapsulation process of the w images in the i-th track, perform the following steps:
  • the image type of the image collected by each camera determines the image type field, and the image type includes at least one of a depth image or a texture image.
  • the encapsulation unit 701 determines the image type field according to the image type of the image captured by the jth camera, it performs the following steps:
  • the image type field When the image type of the image collected by the jth camera is a texture image, set the image type field to a second value; when the image type of the image collected by the jth camera is a depth image, set the image type field to a second value The image type field is set to a third value; when the image type of the image captured by the jth camera is a texture image and a depth image, the image type field is set to a fourth value.
  • the angle of view information corresponding to the i-th track includes a video mosaic layout indication field
  • the w images include texture images and depth images
  • the encapsulation unit 701 according to the w images in the i-th track
  • the video layout indication field is set to the first value; if the texture image and the depth image contained in the i track are coded separately, the video layout indication field is set to the second value.
  • the angle of view information corresponding to the i-th track further includes a camera model field; the encapsulation unit 701 generates the free angle of view corresponding to the i-th track according to the encapsulation process of the w images in the i-th track
  • information data box perform the following steps:
  • the camera model field is set to the first value; when the camera model of the camera that collects the w images belongs to the second model, Set the camera model field to a second numeric value.
  • the viewing angle information corresponding to the i-th track further includes a guard bandwidth field of a texture image and a guard bandwidth field of a depth image
  • the encapsulation unit 701 performs encapsulation according to the w images in the i-th track
  • the angle of view information corresponding to the ith track further includes a second camera identification field and a camera attribute information field
  • the encapsulation unit 701 generates the first For the free viewing angle information data box corresponding to i tracks, perform the following steps:
  • the camera attribute information of the jth camera includes any one or more of the following: the value of the horizontal axis component, the value of the vertical axis component, and the value of the vertical axis component of the position of the jth camera, and the value of the jth camera The value of the horizontal axis component and the vertical axis component of the focal length of the camera, and the resolution width and height of the image captured by the jth camera.
  • the angle of view information corresponding to the i-th track further includes an image information field
  • the encapsulation unit 701 generates the free angle of view corresponding to the i-th track according to the encapsulation process of w images in the i-th track
  • the image information includes one or more of the following: a multiplication factor for depth image downsampling, The offset of the upper left vertex of the depth image relative to the origin of the plane frame, and the offset of the upper left vertex of the texture image relative to the origin of the plane frame.
  • the angle of view information corresponding to the ith track further includes a custom camera parameter field, a custom camera parameter type field, and a custom camera parameter length field;
  • f is an integer greater than or equal to 0 and less than h, and h represents the value of the jth camera Number of custom camera parameters;
  • the generation unit 702 is further configured to: generate a signaling description file corresponding to the immersive media, where the signaling description file includes a free-view camera descriptor corresponding to the immersive media, and the free-view
  • the camera descriptor is used to record the camera attribute information corresponding to the video segment in each track, and the video segment in any track is composed of images encapsulated in any track;
  • the free-view camera descriptor is encapsulated in the In the adaptation set level of the media presentation description file of immersive media, or the signaling description file is encapsulated in the presentation level of the media presentation description file.
  • the data processing apparatus for immersive media further includes a sending unit 703, configured to send the signaling description file to a content consumption device to instruct the content consumption device to The camera attribute information corresponding to the video segment in each track recorded in , select a candidate camera that matches the user's location from the multiple cameras, and send a first resource request to obtain the segmented video from the candidate camera ; In response to the first resource request, select the i-th track from the M tracks according to the viewing angle information in the free viewing angle information data box corresponding to each track in the M tracks and set the i-th track The corresponding free view information data box is sent to the content consumption device.
  • a sending unit 703 configured to send the signaling description file to a content consumption device to instruct the content consumption device to The camera attribute information corresponding to the video segment in each track recorded in , select a candidate camera that matches the user's location from the multiple cameras, and send a first resource request to obtain the segmented video from the candidate camera ; In response to the first resource request, select the i-th track
  • the sending unit 703 is further configured to: send the signaling description file to the content consumption device, so as to instruct the content consumption device to send the second resource request according to the signaling description file and user bandwidth ; In response to the second resource request, send M free view information data boxes corresponding to the M tracks to the content consumption device, so as to instruct the content consumption device to use the free view information data corresponding to each track
  • the viewing angle information and the user's location information in the box, and the free viewing angle information data box corresponding to the i-th track is obtained from the M free viewing angle information data boxes.
  • each step involved in the data processing method for immersive media shown in FIG. 5 may be executed by each unit in the data processing apparatus for immersive media shown in FIG. 7 .
  • S501 described in FIG. 5 may be performed by the packaging unit 702 in the data processing device shown in FIG. 7
  • S502 may be performed by the generating unit 702 in the data processing device shown in FIG. 7 .
  • each unit in the data processing device for immersive media shown in FIG. can be further divided into a plurality of functionally smaller units, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application.
  • the above-mentioned units are divided based on logical functions.
  • the functions of one unit may also be realized by multiple units, or the functions of multiple units may be realized by one unit.
  • the immersive media-based data processing device may also include other units.
  • these functions may also be implemented with the assistance of other units, and may be implemented cooperatively by multiple units.
  • a general-purpose computing device such as a computer including processing elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM) and storage elements.
  • CPU central processing unit
  • RAM random access storage medium
  • ROM read-only storage medium
  • Run a computer program capable of executing the steps involved in the corresponding method as shown in Figure 5, to construct a data processing device for immersive media as shown in Figure 7, and to realize the data of the immersive media in the embodiment of the present application Approach.
  • the computer program can be recorded in, for example, a computer-readable storage medium, loaded into the above-mentioned computing device through the computer-readable storage medium, and run there.
  • an immersive medium is packaged into M tracks.
  • the immersive medium is composed of images taken by N cameras at different viewing angles.
  • One track may include images from at least one camera.
  • the M tracks are Belonging to the same track group, a scene in which an immersive video is encapsulated into multiple tracks is realized; in addition, the content production device generates a free-view information data box for each track, which is indicated by the free-view information data box corresponding to the i-th track.
  • the viewing angle information corresponding to the i-th track such as the specific viewing angle position of the camera, and then when the content consumption device decodes and displays the image in the track according to the viewing angle information corresponding to the i-th track, it can ensure that the decoded and displayed video is consistent with the user's location. The position is more matched, improving the presentation effect of immersive video.
  • FIG. 8 it is a schematic structural diagram of a content consumption device provided by an embodiment of the present application.
  • the content consumption device shown in FIG. 8 may refer to a computer device used by a user of immersive media, and the computer device may be a terminal.
  • the content consumption device shown in FIG. 8 may include a receiver 801 , a processor 802 , a memory 803 and a display/play device 804 . in:
  • the receiver 801 is used to realize decoding and transmission interaction with other devices, and is specifically used to realize the transmission of immersive media between the content production device and the content consumption device. That is, the content consumption device receives through the receiver 901 relevant media resources of the immersive media transmitted by the content production device.
  • the processor 802 or CPU (Central Processing Unit, central processing unit)) is the processing core of the content production device, and the processor 802 is suitable for implementing one or more computer programs, specifically for loading and executing one or more computer programs In this way, the flow of the data processing method for immersive media shown in FIG. 4 is realized.
  • CPU Central Processing Unit, central processing unit
  • the memory 803 is a memory device in the content consumption device for storing computer programs and media resources. It can be understood that the storage 803 here may include a built-in storage medium in the content consumption device, and of course may also include an extended storage medium supported by the content consumption device. It should be noted that the memory 803 may be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory; optionally, it may also be at least one memory located far away from the aforementioned processor.
  • the memory 803 provides storage space for storing the operating system of the content consumption device. Moreover, the storage space is also used to store a computer program, and the computer program is suitable for being invoked and executed by the processor, so as to execute each step of the data processing method for immersive media. In addition, the memory 803 may also be used to store the 3D image of the immersive media processed by the processor, the audio content corresponding to the 3D image, and the information required for rendering the 3D image and audio content.
  • the display/playing device 804 is used to output rendered sound and three-dimensional images.
  • the processor 802 may include a parser 821, a decoder 822, a converter 823, and a renderer 824; wherein:
  • the parser 821 is used to decapsulate the packaged file of the rendered media from the content production device, specifically to decapsulate the media file resource according to the file format of the immersive media, to obtain the audio code stream and the video code stream; Code stream and video code stream are improved to decoder 822;
  • the decoder 822 performs audio decoding on the audio code stream to obtain audio content and provides it to the renderer 824 for audio rendering. In addition, the decoder 822 decodes the video code stream to obtain a 2D image. According to the metadata provided by the media presentation description information, if the metadata indicates that the immersive media has performed the area encapsulation process, the 2D image refers to an encapsulated image; if the metadata indicates that the immersive media has not performed the area encapsulation process, then the planar image is Refers to the projected image.
  • the converter 823 is used to convert a 2D image into a 3D image. If the immersive media has performed the area encapsulation process, the converter 923 will firstly decapsulate the area of the encapsulated image to obtain the projected image. Then the projection image is reconstructed to obtain a 3D image. If the area encapsulation process has not been performed on the rendering medium, the converter 923 will directly reconstruct the projected image to obtain a 3D image.
  • the renderer 824 is used for rendering audio content and 3D images of immersive media. Specifically, the audio content and the 3D image are rendered according to the metadata related to rendering and window in the media presentation description information, and the rendering is completed and delivered to the display/playing device for output.
  • the processor 802 executes each step of the data processing method for immersive media shown in FIG. 4 by calling one or more computer programs in the memory.
  • the memory stores one or more computer programs, and the one or more computer programs are suitable for being loaded by the processor 802 and performing the following steps:
  • the free viewing angle information data box corresponding to the i-th track includes the viewing angle information corresponding to the i-th track, i is an integer greater than or equal to 1 and less than or equal to M; the immersive media is composed of The images taken by N cameras at different viewing angles are composed, the immersive media is packaged into M tracks, and images from at least one camera are packaged in one track, and both N and M are integers greater than 1.
  • the image encapsulated in the i-th track is decoded and displayed according to the viewing angle information in the free viewing angle information data box.
  • an immersive medium is packaged into M tracks.
  • the immersive medium is composed of images taken by N cameras at different viewing angles.
  • One track may include images from at least one camera.
  • the M tracks are Belonging to the same track group, a scene in which an immersive video is encapsulated into multiple tracks is realized; in addition, the content production device generates a free-view information data box for each track, which is indicated by the free-view information data box corresponding to the i-th track.
  • the viewing angle information corresponding to the i-th track such as the specific viewing angle position of the camera, and then when the content consumption device decodes and displays the image in the track according to the viewing angle information corresponding to the i-th track, it can ensure that the decoded and displayed video is consistent with the user's location. The position is more matched, improving the presentation effect of immersive video.
  • the content production device may refer to a computer device used by an immersive media provider, and the computer device may be a terminal or a server.
  • the content production device may include a capture device 901 , a processor 902 , a memory 903 and a transmitter 904 . in:
  • the capture device 901 is used to collect audio-visual scenes of the real world to obtain original data of immersive media (including audio content and video content that are synchronized in time and space).
  • the capture device 801 may include, but not limited to: audio equipment, camera equipment, and sensor equipment.
  • the audio device may include an audio sensor, a microphone, and the like.
  • the camera device may include a common camera, a stereo camera, a light field camera, and the like.
  • Sensing devices may include laser devices, radar devices, and the like.
  • the processor 902 (or called CPU (Central Processing Unit, central processing unit)) is the processing core of the content production equipment, and the processor 902 is suitable for implementing one or more computer programs, and is specifically suitable for loading and executing one or more computer programs.
  • the program thus realizes the flow of the data processing method for immersive media shown in FIG. 4 .
  • the storage 903 is a storage device in the content production device, and is used to store programs and media resources. It can be understood that the storage 903 here may include a built-in storage medium in the content production device, and of course may also include an extended storage medium supported by the content production device. It should be noted that the memory may be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory; optionally, it may also be at least one memory located away from the aforementioned processor. The memory provides storage space for storing the operating system of the content production device. Moreover, the storage space is also used to store computer programs, the computer programs include program instructions, and the program instructions are suitable to be invoked and executed by the processor, so as to execute each step of the data processing method for immersive media. In addition, the memory 903 can also be used to store the immersive media file processed by the processor, and the immersive media file includes media file resources and media presentation description information.
  • the transmitter 904 is used to realize the transmission interaction between the content production device and other devices, and is specifically used to realize the transmission of immersive media between the content production device and the content consumption device. That is, the content production device transmits relevant media resources of the immersive media to the content consumption device through the transmitter 904 .
  • the processor 902 may include a converter 921 , an encoder 922 and an encapsulator 923 . in:
  • the converter 921 is used to perform a series of conversion processes on the captured video content, so as to make the video content suitable for video encoding of immersive media.
  • the conversion processing may include: splicing and projection, and optionally, the conversion processing may also include region encapsulation.
  • the converter 921 can convert the captured 3D video content into a 2D image, and provide it to the encoder for video encoding.
  • the encoder 922 is configured to perform audio encoding on the captured audio content to form an audio stream of immersive media. It is also used to perform video coding on the 2D image converted by the converter 921 to obtain a video code stream.
  • the encapsulator 923 is used to encapsulate the audio code stream and the video code stream in a file container according to the file format of the immersive media (such as ISOBMFF) to form a media file resource of the immersive media.
  • the media file resource can be a media file or a media segment to form the immersive media.
  • the media file of the immersive media and record the metadata of the media file resource of the immersive media by using the media presentation description information according to the file format requirements of the immersive media.
  • the package file of the immersive media processed by the packager will be stored in the memory, and provided to the content consumption device on demand for presentation of the immersive media.
  • the processor 902 executes each step of the data processing method for immersive media shown in FIG. 5 by calling one or more instructions in the memory.
  • the memory 803 stores one or more computer programs, which are suitable for being loaded and executed by the processor 902 but are not limited to the following steps:
  • the immersive media is composed of images taken by N cameras at different viewing angles, one track encapsulates images from at least one camera, and the M tracks belong to the same track group, N and M are both integers greater than or equal to 1;
  • the free view angle information data box includes the view angle information corresponding to the i-th track; 1 ⁇ i ⁇ M , w ⁇ 1.
  • an immersive medium is packaged into M tracks.
  • the immersive medium is composed of images taken by N cameras at different viewing angles.
  • One track may include images from at least one camera.
  • the M tracks are Belonging to the same track group, a scene in which an immersive video is encapsulated into multiple tracks is realized; in addition, the content production device generates a free-view information data box for each track, which is indicated by the free-view information data box corresponding to the i-th track.
  • the viewing angle information corresponding to the i-th track such as the specific viewing angle position of the camera, and then when the content consumption device decodes and displays the image in the track according to the viewing angle information corresponding to the i-th track, it can ensure that the decoded and displayed video is consistent with the user's location. The position is more matched, improving the presentation effect of immersive video.
  • an embodiment of the present application further provides a storage medium, where the storage medium is used to store a computer program, and the computer program is used to execute the method provided in the foregoing embodiments.
  • the embodiment of the present application also provides a computer program product including instructions, which, when run on a computer, causes the computer to execute the method provided in the foregoing embodiments.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Library & Information Science (AREA)
  • Television Signal Processing For Recording (AREA)
  • Testing, Inspecting, Measuring Of Stereoscopic Televisions And Televisions (AREA)
  • Signal Processing For Digital Recording And Reproducing (AREA)
  • Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

本申请实施例公开了一种沉浸媒体的数据处理方法、装置、设备及存储介质,其中沉浸媒体是由处于不同视角的N个相机拍摄的图像组成,沉浸媒体被封装到M个轨道,一个轨道中封装来自至少一个相机的图像,N和M均为大于1的整数;包括:获取沉浸媒体的第i个轨道对应的自由视角信息数据盒,自由视角信息数据盒包括第i个轨道对应的视角信息(S401),i为大于或等于1且小于或等于M的整数;根据自由视角信息数据盒中的视角信息对第i个轨道内封装的图像进行解码(S402)。采用本申请实施例可以使得内容消费设备根据各个轨道中的视角信息以及用户当前位置选择合适的图像进行解码消费。

Description

沉浸媒体的数据处理方法、装置、相关设备及存储介质
本申请要求于2021年06月11日提交中国专利局、申请号为202110659190.8、申请名称为“沉浸媒体的数据处理方法、装置、相关设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及音视频领域,尤其涉及媒体的数据处理。
背景技术
沉浸媒体是指能为消费者带来沉浸式体验的媒体内容,沉浸媒体又可以称为自由视角视频,自由视角视频通常是由相机阵列从多个角度对同一个三维场景进行拍摄,得到不同视角的深度图像和/或纹理图像,这些深度图像和/或纹理图像组成了自由视角视频。内容消费设备可以根据用户当前所在位置以及各个图像来源的相机视角,选择解码某些图像进行消费。
目前在自由视角视频制作过程中一般是采用大规模图集信息数据盒指示自由视角视频相关的参数信息(比如相机采集的深度图以及纹理图的分辨力宽度与高度、每个视角对应的相机标识符等等),从而省略图集轨道中的其余图集信息。
发明内容
本申请实施例提供了一种沉浸媒体的数据处理方法、装置、设备及存储介质,可以将一沉浸媒体的不同视角相机采集的图像封装到多个不同轨道,并采用每个轨道对应的自由视角信息数据盒来指示每个轨道中图像来源相机的视角信息,以便于内容消费设备根据各个轨道中的视角信息以及用户当前位置选择合适的图像进行解码消费。
一方面,本申请实施例提供了一种沉浸媒体的数据处理方法,该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成,所述沉浸媒体被封装到M个轨道,一个轨道中封装来自至少一个相机的图像,N和M均为大于1的整数,数据处理方法包括:
获取第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒子,包括所述第i个轨道对应的视角信息,i为大于等于1且小于等于M的整数;
根据所述自由视角信息数据盒中的视角信息对所述第i个轨道内封装的图像进行解码。
一方面,本申请实施例提供了另一种沉浸媒体的处理方法,包括:
将沉浸媒体封装到M个轨道中,所述沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中封装来自至少一个相机的图像,所述M个轨道属于同一个轨道组,N和M均为大于或等于1的整数;
根据第i个轨道中w个图像的封装过程,生成第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒包括所述第i个轨道对应的视角信息;1≤i≤M,w≥1。
一方面,本申请实施例提供了一种沉浸媒体的数据处理装置,包括:
获取单元,用于获取沉浸媒体的第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒子,包括所述第i个轨道对应的视角信息,沉浸媒体是由处于不同视角的N个相机拍摄的图像组成,所述沉浸媒体被封装到M个轨道,一个轨道中封装来自至少一个相机的图像,N和M均为大于1的整数;i为大于或等于1且小于或等于M的整数;
解码单元,用于根据所述自由视角信息数据盒中的视角信息对所述第i个轨道内封装的图像进行解码。
一方面,本申请实施例提供了另一种沉浸媒体的数据处理装置,包括:
封装单元,用于将沉浸媒体封装到M个轨道中,所述沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中封装来自至少一个相机的图像,所述M个轨道属于同一个轨道组,N和M均为大于或等于1的整数;
生成单元,用于根据第i个轨道中w个图像的封装过程,生成第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒包括所述第i个轨道对应的视角信息;1≤i≤M,w≥1。
一方面,本申请实施例提供了一种内容消费设备,包括:
处理器,适于实现一条或多条计算机程序;以及
计算机存储介质,所述计算机存储介质存储有一条或多条计算机程序,所述一条或计算机程序程序适于由处理器加载并执行:
获取沉浸媒体的第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒子,包括所述第i个轨道对应的视角信息;该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成,所述沉浸媒体被封装到M个轨道,一个轨道中封装来自至少一个相机的图像,N和M均为大于1的整数,i为大于或等于1且小于或等于M的整数;
根据所述自由视角信息数据盒中的视角信息对所述第i个轨道内封装的图像进行解码。
一方面,本申请实施例提供了一种内容制作设备,包括:
处理器,适于实现一条或多条计算机程序;以及
计算机存储介质,所述计算机存储介质存储有一条或多条计算机程序,所述一条或计算机程序程序适于由处理器加载并执行:
将沉浸媒体封装到M个轨道到M个轨道中,所述沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中封装来自至少一个相机的图像,所述M个轨道属于同一个轨道组,N和M均为大于或等于1的整数;
根据第i个轨道中w个图像的封装过程,生成第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒包括所述第i个轨道对应的视角信息;1≤i≤M,w≥1。
又一方面,本申请实施例提供一种存储介质,所述存储介质用于存储计算机程序,所述计算机程序用于执行以上方面的方法。
又一方面,本申请实施例提供了一种包括指令的计算机程序产品,当其在计算机上运行时,使得所述计算机执行以上方面的方法。
本申请实施例中将一沉浸媒体封装至M个轨道,该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中可以包括来自至少一个相机的图像,这M个轨道是属于同一个轨道组的,实现了一个沉浸视频封装到多个轨道的场景;另外,内容制作设备为每个轨道生成一个自由视角信息数据盒,通过第i个轨道对应的自由视角信息数据盒指示第i个轨道对应的视角信息,比如相机所在的具体视角位置,进而内容消费设备在根据第i个轨道对应的视角信息进行该轨道内图像的解码时,可以保证解码的视频与用户所在位置更加 匹配,提高沉浸视频的呈现效果。
附图说明
图1a是本申请实施例提供的一种用户消费3DoF沉浸媒体的示意图;
图1b是本申请实施例提供的一种用户消费3DoF+沉浸媒体的示意图;
图1c是本申请实施例提供的一种用户消费6DoF沉浸视频的示意图;
图2a是本申请实施例提供的一种沉浸媒体系统的架构图;
图2b是本申请实施例提供的一种沉浸媒体的传输方案的示意图;
图3a是本申请实施例提供的一种视频编码基本框图;
图3b是本申请实施例通过的一种输入图像划分的示意图;
图4是本申请实施例提供的一种沉浸媒体的数据处理方法的流程示意图;
图5是本申请实施例提供的另一种沉浸媒体的数据处理方法的流程示意图;
图6是本申请实施例提供的一种沉浸媒体的数据处理装置的结构示意图;
图7是本申请实施例提供的另一种沉浸媒体的数据处理装置的结构示意图;
图8是本申请实施例提供的一种内容消费设备的结构示意图;
图9是本申请实施例提供的一种内容制作设备的结构示意图。
具体实施方式
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述。
本申请实施例涉及到沉浸媒体的数据处理技术,所谓沉浸媒体是指能够提供沉浸式的媒体内容,使沉浸于该媒体内容中的用户能够获得现实世界中视觉、听觉等感官体验的媒体文件。具体的,沉浸媒体可以是三自由度(3Degree of Freedom,3DoF)沉浸媒体,3DoF+沉浸媒体或者六自由度(6Degree of Freedom,6DoF)沉浸媒体。
参见图1a,为本申请实施例提供的一种用户消费3DoF沉浸媒体的示意图,图1a所示的3DoF沉浸媒体是指用户在一个三维空间的中心点固定,用户头部沿着X轴、Y轴和Z轴旋转来观看沉浸媒体的媒体内容提供的画面。参见图1b,为本申请实施例提供的一种用户消费3DoF+沉浸媒体的示意图,3DoF+是指当沉浸媒体提供的虚拟场景具有一定的深度信息,用户头部可以基于3DoF在一个有限的空间内移动来观看媒体内容提供的画面。参见图1c,为本申请实施例提供的一种用户消费6DoF沉浸视频的示意图,6DoF分为窗口6DoF、全方向6DoF和6DoF,其中,窗口6DoF是指用户在X轴、Y轴的旋转移动受限,以及在Z轴的平移受限;例如,用户不能够看到窗户框架外的景象,以及用户无法穿过窗户。全方向6DoF是指用户在X轴、Y轴和Z轴的旋转移动受限,例如,用户在受限的移动区域中不能自由的穿过三维的360度VR内容。6DoF是指用户可以沿着X轴、Y轴、Z轴自由平移,例如,用户可以在三维的360度VR内容中自由的走动。简单来讲,6DoF沉浸视频不仅可以允许用户沿着X轴、Y轴以及Z轴旋转消费媒体内容,还可以沿着X轴、Y轴以及Z轴自由运动来消费媒体内容。
沉浸媒体内容包括以各种形式在三维(3-Dimension,3D)空间中表示的视频内容,例如以球面形式表示的三维视频内容。具体地,沉浸媒体内容可以是VR(Virtual Reality,虚 拟现实)视频内容、多视角视频内容、全景视频内容、球面视频内容或360度视频内容;所以,沉浸媒体又可称为VR视频、自由视角视频、全景视频、球面视频或360度视频。另外,沉浸媒体内容还包括与三维空间中表示的视频内容相同步的音频内容。
参见图2a,为本申请实施例提供的一种沉浸媒体系统的架构图。在图2a所示的沉浸媒体系统中包括内容制作设备和内容消费设备,内容制作设备可以指沉浸媒体的提供者(例如沉浸内容的内容制作者)所使用的计算机设备,该计算机设备可以是终端,比如智能手机、平板电脑、笔记本电脑、台式计算机、智能音箱、智能手表、智能车载等;该计算机设备也可以是服务器,比如独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、CDN、以及大数据和人工智能平台等基础云计算服务的云服务器。
内容消费设备可以指沉浸媒体的使用者(例如用户)所使用的计算机设备,该计算机设备可以是终端,比如个人计算机、智能移动设备如智能手机、VR设备(例如VR头盔、VR眼镜等)。沉浸媒体的数据处理过程包括在内容制作设备侧的数据处理过程以及在内容消费设备侧的数据处理过程。
在内容制作设备侧的数据处理过程主要包括:(1)沉浸媒体的媒体内容的获取与制作过程;(2)沉浸媒体的编码及封装的过程。在内容消费侧的数据处理过程主要包括:(3)沉浸媒体的解封装及解码的过程;(4)沉浸媒体的渲染过程。另外,内容制作设备与内容消费设备之间涉及沉浸媒体的传输过程,该传输过程可以基于各种传输协议来进行,此处的传输协议可以包括但不限于:DASH(Dynamic Adaptive Streaming over HTTP,动态自适应流媒体传输)协议、HLS(HTTP Live Streaming,动态码率自适应传输)协议、SMTP(Smart Media TransportProtocol,智能媒体传输协议)、TCP(Transmission Control Protocol,传输控制协议)等。
参见图2b,为本申请实施例提供的一种沉浸媒体的传输方案的示意图。如图2b所示,为了解决沉浸媒体自身数据量过大带来的传输带宽负荷问题,在沉浸媒体的处理过程中,通常选择将原始的沉浸媒体在空间上切分为多个分块视频后,分别编码后进行封装,再传输给客户端消费。下面结合图2b,分别对沉浸媒体的数据处理过程进行详细介绍。首先介绍内容制作设备侧的数据处理过程:
(1)沉浸媒体的媒体内容的获取与制作过程:
S1:沉浸媒体的媒体内容的获取过程。
沉浸媒体的媒体内容是通过捕获设备采集现实世界的声音-视觉场景获得的。在一个实施例中,捕获设备可以是指设于内容制作设备中的硬件组件,例如捕获设备是指终端的麦克风、摄像头以及传感器等等。在其他实施例中,该捕获设备也可以是独立与内容制作设备但与内容制作设备相连接的硬件装置,例如与服务器相连接的摄像头。该捕获设备可以包括但不限于:音频设备、摄像设备及传感设备。其中,音频设备可以包括音频传感器、麦克风等。摄像设备可以包括普通摄像头、立体摄像头、光场摄像头等。传感设备可以包括激光设备、雷达设备等。捕获设备的数据可以为多个,这些捕获设备可以被部署在现实 空间中的一些特定视角以同时捕获该空间内不同视角的音频内容以及视频内容,捕获的音频内容和视频内容在时间和空间上均保持同步。举例来说,3DoF沉浸内容的媒体内容是由一组摄像机或一个带有多个摄像头和传感器的摄像设备录制的,6DoF沉浸媒体的媒体内容主要由相机阵列拍摄得到的点云、光场等形式的内容制作而成。
S2:沉浸媒体的媒体内容的制作过程。
捕获到的音频内容本身就是适合被执行沉浸媒体的音频编码的内容,因此无需对捕获到的音频内容进行其他处理。而捕获到的视频内容需要进行一系列制作流程后才可以称为适合被执行沉浸媒体的视频编码的内容,该制作流程具体可以包括:
①拼接,由于捕获到的沉浸媒体的视频内容是捕获设备在不同视角下拍摄得到的,拼接就是指对这些各个视角拍摄的视频内容拼接成一个完整的、能够反映现实空间360度视觉全景的视频,即拼接后的视频是一个在三维空间表示的全景视频。
②投影,投影就是指将拼接形成的一个三维视频映射到一个二维(2-Dimension,2D)图像上的过程,投影形成的2D图像称为投影图像;投影的方式可包括但不限于:经纬图投影、正六面体投影。
需要说明的是,由于采用捕获设备只能捕获到全景视频,这样的视频经内容制作设备处理并传输至内容消费设备进行相应的数据处理后,内容消费设备侧的用户只能通过一些特定动作(如头部旋转)来观看360度的视频信息,而执行非特定动作(如移动头部)并不能获得相应的视频变化,VR体验不佳,因此需要额外提供与全景视频相匹配的深度信息,来使用户获得更优的沉浸度和更佳的VR体验,这就涉及多种制作技术,常见的制作技术包括6DoF制作技术、3DoF制作技术以及3DoF+制作技术。
采用6DoF制作技术和3DoF+制作技术得到的沉浸媒体可以包括自由视角视频。自由视角视频作为一种常见的3DoF+和6DoF沉浸式媒体,其是由多个相机采集的、包含不同视角的,支持用户3DoF+或6DoF交互的沉浸式媒体视频。具体的,3DoF+沉浸式媒体是由一组相机或一个带有多个摄像头和传感器的相机录制而成,相机通常可以获取在设备中心周围所有方向的内容。6DoF沉浸式媒体主要由相机阵列拍摄得到的点云、光场等形式的内容制作而成。
(3)沉浸媒体的媒体内容的编码过程。
投影图像可以被直接进行编码,也可以对投影图像进行区域封装之后再进行编码。参见图3a,为本申请实施例提供的一种视频编码基本框图。现代主流视频编码技术,以国际视频编码标准HEVC(High Efficiency Video Coding),国际视频编码标准VVC(Versatile Video Coding),以及中国国家视频编码标准AVS(Audio Video Coding Standard)为例,采用了混合编码框架,对输入的原始视频信号,进行了如下一系列的操作和处理:
1)块划分结构(block partition structure):根据处理单元的大小将输入图像划分成若干个不重叠的处理单元,对每个处理单元进行类似的压缩操作。这个处理单元被称为编码树单元(Coding Tree Unit,CTU),或者最大编码单元(Largest Coding Unit,LCU)。CTU可以继续进行更加精细的划分,得到一个或多个基本编码的单元,称为编码单元(Coding Unit,CU)。每个CU是一个编码缓解中最基本的元素。参见图3b,为本申请实施例通过的一种输 入图像划分的示意图。以下描述的是对每一个CU可能采用的各种编码方式。
2)预测编码(Predictive Coding):包括帧内预测(Intra(picture)Prediction)和帧间预测(Inter(picture)Prediction)。沉浸媒体的原始视频信号经过选定的已重建视频信号的预测后,得到残差视频信号。内容制作设备需要为当前CU决定在众多可能的预测编码模式中,选择最合适的一种,并告知内容消费设备。其中,帧内预测所预测的信号来自于同一个图像内已经过编码重建过的区域,帧间预测预所预测的信号来自已经编码过的,不同于当前图像的其他图像(称之为参考图像)。
3)变换编码及量化(Transform&Quantization):残差视频信号经过离散傅里叶变换(Discrete Fourier Transform,DFT),离散余弦变换(Discrete Cosine Transform,DCT)等变换操作,将信号转换到变换域中,称之为变换系数。在变换域中的信号,进一步的进行有损的量化操作,丢失掉一定的信息,使得量化后的信号有利于压缩表达。在一些视频编码标准中,可能有多于一种变换方式可以选择,因此,内容制作设备也需要为当前编码CU选择其中的一种变换,并告知内容播放设备。量化的精细程度通常由量化参数(Quantization Parameter,QP)来决定,QP取值较大,表示更大取值范围的系数将被量化为同一个输出,因此通常会带来更大的失真,及较低的码率;相反,QP取值较小,表示较小取值范围的系数将被量化为同一个输出,因此通常会带来较小的失真,同时对应较高的码率。
4)熵编码(Entropy Coding)或统计编码:量化后的变换域信号,将根据各个值出现的频率,进行统计压缩编码,最后输出二值化(0或者1)的压缩码流。同时,编码产生其他信息,例如选择的模式,运动矢量等,也需要进行熵编码以降低码率。统计编码是一种无损编码方式,可以有效的降低表达同样的信号所需要的码率。常见的统计编码方式有变长编码(VLC,Variable Length Coding)或者基于上下文的二值化算术编码(CABAC,Content Adaptive Binary Arithmetic Coding)。
5)环路滤波(Loop Filtering):已经编码过的图像,经过反量化,反变换及预测补偿的操作(上述2~4的反向操作),可获得重建的解码图像。重建图像与原始图像相比,由于存在量化的影响,部分信息与原始图像有所不同,产生失真(Distortion)。对重建图像进行滤波操作,例如去块效应滤波(deblocking),取样自适应偏移(Sample Adaptive Offset,SAO)滤波器或者自适应环路滤波器(Adaptive Loop Filter,ALF)等,可以有效的降低量化所产生的失真程度。由于这些经过滤波后的重建图像,将作为后续编码图像的参考,用于对将来的信号进行预测,所以上述的滤波操作也被称为环路滤波,及在编码环路内的滤波操作。
此处需要说明的是,如果采用6DoF(Six Degrees of Freedom,六自由度)制作技术(用户可以在模拟的场景中较自由的移动时,称为6DoF),在视频编码过程中需要采用特定的编码方式(如点云编码)进行编码。
(4)沉浸媒体的封装过程。
将音频码流和视频码流按照沉浸媒体的文件格式(如ISOBMFF(ISO Base Media File Format,ISO基媒体文件格式))封装到文件容器(轨道)中形成沉浸媒体的媒体资源文件,该媒体资源文件可以是媒体文件或者媒体片段形成的沉浸媒体的媒体文件,并按照沉浸媒体的文件格式要求采用媒体呈现描述信息(Media presentation description,MPD)记录该沉 浸媒体的媒体文件资源的元数据,此处的元数据是对于沉浸媒体的呈现有关的信息的总称,该元数据可以包括对媒体内容的描述信息、对视窗的描述信息以及对媒体内容呈现相关的信令信息等等。如图2a所示,内容制作设备会存储经过数据处理过程之后形成的媒体呈现描述信息和媒体文件资源。
下面介绍内容消费设备侧的数据处理过程:
(1)沉浸媒体的解封以及解码的过程:
内容消费设备可以通过内容制作设备的推荐或者按照内容消费设备侧用户需求自适应动态从内容制作设备获得沉浸媒体的媒体文件资源和相应媒体呈现描述信息,例如内容消费设备可以根据用户的头部/眼睛/身体的跟踪信息确定用户的朝向和位置,再基于确定的朝向和位置动态向内容制作设备请求获得相应的媒体文件资源。媒体文件资源和媒体呈现描述信息通过传输机制(如DASH、SMT)由内容制作设备传输给内容消费设备。内容消费设备的解封装过程与内容制作设备的封装过程是相逆的,内容消费设备按照沉浸媒体的文件格式要求对获取到的媒体文件资源进行解封装,得到音频码流和视频码流。内容消费设备的解码过程与内容制作设备的编码过程是相逆的,内容消费设备对音频码流进行音频解码,还原出音频内容,以及内容消费设备对视频码流进行解码,得到视频内容。其中,内容消费设备对视频码流的解码过程可以包括如下:①对视频码流进行解码,得到平面的投影图像。②根据媒体呈现描述信息将投影图像进行重建处理以转换为3D图像,此处的重建处理是指二维的投影图像重新投影至3D空间中的处理。
根据上述编码过程可以看出,在内容消费设备侧,对于每一个CU,内容消费设备获得压缩码流后,先进行熵解码,获得各种模式信息以及量化后的变换系数。各个系数经过反复量化以及变换,得到残差信号。另一方面,根据已知的编码模式信息,可获得该CU对应的预测信号,两者相加之后,即可得到重建信号。最后解码图像的重建值,需要经过环路滤波的操作,产生最终的输出信号。
(2)沉浸媒体的渲染过程。
内容消费设备根据媒体沉陷描述信息中与渲染、视窗相关的元数据对音频解码得到的音频内容以及视频解码得到的3D图像进行渲染,渲染完成即实现了对该3D图像的播放输出。特别地,如果采用3DoF和3DoF+的制作技术,内容消费设备主要基于当前视点、视差、深度信息等对3D图像进行渲染,如果采用6DoF的制作技术,内容消费设备主要基于当前视点对视窗内的3D图像进行渲染。其中,视点指用户的观看位置点,视差是指用户的双目产生的视线差或由于运动产生的视线差,视窗是指观看区域。
上述描述的沉浸媒体系统支持数据盒(Box),数据盒是指包括元数据的数据块或对象,即数据盒子中包括了相应媒体内容的元数据。由上述沉浸媒体的数据处理过程可知,在对沉浸媒体进行编码后,需要对编码后的沉浸媒体进行封装并传输给用户。本申请实施例中沉浸媒体主要指自由视角视频,现有技术中考虑到自由视角视频在制作过程中,图集信息仅有相机参数即可获取,纹理图像和深度图像在平面帧中的位置也较为固定时,可使用大规模图集信息数据盒指示相关的参数信息,从而省略图集轨道中的其余图集信息。
具体实现中,大规模图集信息数据盒的语法可参见下述代码段1所示:
Figure PCTCN2022080257-appb-000001
上述代码段1所示语法的语义如下:camera_count表示采集沉浸媒体的所有相机的个数;padding_size_depth表示在对深度图像进行编码时采用的保护带宽度;padding_size_texture表示对纹理图像进行编码时采用的保护带宽度;camera_id表示处于一个视角相机的相机标识符,camera_resolution_x表示一个相机采集的纹理图像、深度图像的分辨率宽度,camera_resolution_y表示一个相机采集的纹理图像、深度图像的分辨力高度;depth_downsample_factor表示深度图像降采样的倍数因子,深度图像的实际分辨率宽度与高度为相机采集分辨率宽度与高度的1/2depth_downsample_factor;depth_vetex_x表示深度图像左上顶点相对于平面帧原点(平面帧的左上顶点)偏移的横轴分量,depth_vetex_y表示深度图像左上顶点相对于平面帧原点偏移的纵轴分量;texture_vetex_x表示纹理图像的左上顶点相对于平面帧原点偏移的横轴分量,texture_vetex_y表示纹理图像的左上顶点相对于平面帧原点偏移的纵轴分量;camera_para_length表示容积视频重构时所需的相机参数的长度,以字节为单位;camera_parameter表示容积视频重构时所需的相机参数。
从上面的大规模图集信息数据盒的语法中可以看出,虽然大规模图集信息数据盒中指示了自由视角视频帧中纹理图像和深度图像的布局信息,并给出了相关的相机参数比如camera_resolution_y以及camera_resolution_x等,但是上述只考虑了自由视角视频封装在单轨道的场景,没有考虑到自由视角视频被封装到多轨道的场景。并且,上述大规模图集信息数据盒中指示了自由视角视频内纹理图和深度图的排布信息以及相关的相机参数,但是只考虑了将自由视角视频封装到单轨道的情况,并未考虑多轨道封装的场景,另外大规模 图集信息数据盒中指示的相机参数,无法作为内容消费设备选择不同视角的图像进行解码消费的依据,也就是说根据上述大规模图集信息数据盒中记载的相机参数,内容消费设备无法知道哪个图像是适合当前用户位置信息的,从而给内容消费设备解码带来不便。
基于此,本申请实施例提供了一种沉浸媒体的数据处理方案,在该数据处理方案中将沉浸媒体封装到M个轨道中,该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,这M个轨道是属于同一个轨道组的,实现了一个沉浸视频封装到多个轨道的场景;另外,内容制作设备为每个轨道生成一个自由视角信息数据盒,通过第i个轨道对应的自由视角信息数据盒指示第i个轨道对应的视角信息,比如相机所在的具体视角位置,进而内容消费设备在根据第i个轨道对应的视角信息进行该轨道内图像的解码显示时,可以保证解码显示的视频与用户所在位置更加匹配,提高沉浸视频的呈现效果。
基于上述描述,本申请实施例提供了一种沉浸媒体的数据处理方法,参见图4,为本申请实施例提供的一种沉浸媒体的数据处理方法的流程示意图。图4所述的数据处理方法可由内容消费设备执行,具体可由内容消费设备的处理器执行。图4所示的数据处理方法可包括如下步骤:
S401、获取沉浸媒体的第i个轨道对应的自由视角信息数据盒,该自由视角信息数据盒包括第i个轨道对应的视角信息。
其中,沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,沉浸媒体被封装到M个轨道中,一个轨道中封装来自至少一个相机拍摄的图像,这M个轨道属于同一个轨道组。第i个轨道是指M个轨道中被选择的一个,后面具体介绍如何从M个轨道中选择第i个轨道。
封装了沉浸视频的这M个轨道通过每个轨道内封装的自由视角轨道组信息关联的。下面以第i个轨道为例,第i个轨道内封装有第i个轨道对应的自由视角轨道组信息,该自由视角轨道组信息用于指示第i个轨道与封装了沉浸视频的其他轨道属于同一个轨道组。第i个轨道内封装的自由视角轨道组信息可以通过扩展轨道组数据盒得到,第i个轨道内封装的自由视角轨道组信息的语法可以表示为如下代码段2所示:
Figure PCTCN2022080257-appb-000002
在上述代码段2中,自由视角轨道组信息是通过扩展轨道组数据盒得到,以“a3fg”轨道组类型标识。在所有包含“afvg”类型TrackGroupTypeBox(轨道组数据盒)的轨道中,组ID相同的轨道属于同一个轨道组。
在代码段2中,假设第i个轨道内封装了前述N个相机中k个相机采集的图像,任一相机 采集的图像可以是纹理图像或深度图像中的至少一种。代码段2中,自由视角轨道组信息的语法的含义可以如下:camera_count表示第i个轨道组内包含纹理图像或深度图像来源的相机的数目,例如camera_count=k,表示第i个轨道内的封装了来自k个相机的纹理图像和/或深度图像。一个相机对应一个标识信息,每个标识信息存储在自由视角轨道组信息的第一相机标识字段camera_id中,应当理解的,一个相机对应一个标识信息,第i个轨道内包括k个相机,因此,自由轨道组信息中包括k个第一相机标识字段camera_id。
自由轨道组信息中还包括图像类型字段depth_texture_type,该图像类型字段用于指示第j个相机采集图像所属的图像类型,j大于0且小于k,也就是说一个图像类型字段depth_texture_type用于指示一个相机采集图像所属的图像类型中,由于第i个轨道中的纹理图像或者深度图像来源于k个相机,因此,自由轨道组信息中包括k个图像类型字段。
指示第j个相机采集的图像所属图像类型的图像类型字段depth_texture_type具体可参见下述表1所示:
表1
depth_texture_type 含义
0 保留
1 表明包含对应相机拍摄的纹理图像
2 表明包含对应相机拍摄的深度图像
3 表明包含对应相机拍摄的纹理图像和深度图像
具体来讲,当图像类型字段depth_texture_type为1时,表明第j个相机采集的图像所属图像类型为纹理图像;当图像类型字段depth_texture_type为2时,表明第j个相机采集的图像所属图像类型为深度图像;当图像类型字段depth_texture_type为3时指示第j个相机采集的图像所属图像类型为纹理图像和深度图像。
由S401可知,第i个轨道对应的自由视角信息数据盒AvsFreeViewInfoBox中包括了第i个轨道对应的视角信息,参见下代码段3,为第i个轨道对应的自由视角信息数据盒的语法表示,下面结合代码段3具体介绍第i个轨道对应的自由视角信息数据盒以及自由视角信息数据盒中包括的视角信息。其中,代码段3具体如下:
Figure PCTCN2022080257-appb-000003
Figure PCTCN2022080257-appb-000004
①第i个轨道对应的视角信息可以包括视频拼接布局指示字段stitching_layout,该视频拼接布局字段主要用于指示第i个轨道内包括的纹理图像和深度图像是否拼接编码;具体地,当视频拼接布局指示字段stitching_layout为第一数据值时,指示第i个轨道内包括的纹理图像和深度图像是拼接编码的;当视频拼接布局指示字段stitching_layout为第二数值时,指示第i个轨道内包括的纹理图像和深度图像是分别编码的。举例来说,假设沉浸媒体是指6DoF视频,第一数值为0,第二数值为1,则在第i个轨道内视频拼接布局指示字段stitching_layout可以参见下述表2所示:
表2
stitching_layout 6DoF视频拼接布局
0 深度图和纹理图拼接编码
1 深度图和纹理图分别编码
其他 保留
②第i个轨道对应的视角信息还包括相机模型字段camera_model,该字段用于指示第i个轨道内k个相机的相机模型,由前述可知,k表示第i个轨道内深度图像和/或纹理图像来源的相机数量,在第i个轨道对应的视角信息中还可以包括相机数量字段camera_count,该相 机数量字段用于存储第i个轨道内深度图像和/或纹理图像来源的相机数量,假设表示为k。
具体地,当相机模型字段camera_model为第三数值时,指示第j个相机所属相机模型为第一模型;当相机模型字段camera_model为第四数值时,指示第j个相机所属相机模型为第二模型。其中,第一模型可以指针孔模型,第二模型可以指鱼眼模型,假设第三数值为0,第四数值为1,沉浸媒体为6DoF视频,第i个轨道对应的视角信息中相机模型字段camera_model可参见下述表3所示:
表3
Camera model 6DoF视频相机模型
0 针孔模型
1 鱼眼模型
其他 保留
③第i个轨道对应的视角信息还包括纹理图像的保护带宽度字段texture_padding_size,以及深度图像的保护带宽度字段depth_padding_size。纹理图像的保护带宽度字段用于存储第i个轨道内对纹理图像编码时采用的保护带宽度,深度图像的保护带宽度用于存储第i个轨道内对深度图像编码时采用的保护带宽度。
④第i个轨道对应的视角信息还包括第二相机标识字段camera_id,第二相机标识字段camera_id用于存储第i个轨道内第j个相机的标识信息,由前述描述可知,第i个轨道内包括的纹理图像和深度图像来源这k个相机,j的取值是大于等于0,且小于k。也就是说一个第二相机标识字段camera_id存储k个相机中任意一个相机的标识信息,因此,需要k个相机标识字段来存储k个相机的标识信息。需要说明的是,这里的第二相机标识字段camera_id和前述的自由视角轨道组信息中的第一相机标识字段camera_id的作用相同,均是用于存储第i个轨道内第j个相机的标识信息。
⑤第i个轨道对应的视角信息还包括相机属性信息字段,相机属性信息字段用于存储第j个相机的相机属性信息,第j个相机的相机属性信息可以包括第j个相机位置的横轴分量取值、纵轴分量取值以及竖轴分量取值,第j个相机焦距的横轴分量取值和纵轴分量取值,以及第j个相机采集图像的分辨率宽度与高度。因此,相机属性信息字段可具体包括:1)相机位置的横轴分量字段camera_pos_x,用于存储第j个相机位置的横轴分量取值(也称为x分量取值);2)相机位置的纵轴分量字段camera_pos_y,用于存储第j个相机位置的纵轴分量取值(也称为y分量取值);3)相机位置的竖轴分量字段camera_pos_z,用于存储第j个相机位置的竖轴分量取值(也称为z分量取值);4)相机焦距的横轴分量字段focal_length_x,用于存储第j个相机焦距的横轴分量取值(也称为x分量取值);5)相机焦距的纵轴分量字段focal_length_y,用于存储第j个相机焦距的纵轴分量取值(也称为y分量取值);6)相机采集图像的分辨率宽度字段camera_resolution_x,用于存储第j个相机采集图像的分辨率宽度;7)相机采集图像的分辨率宽度字段camera_resolution_y,用于存储第j个相机采集图像的分辨率高度。
需要说明的是,由于第i个轨道内的纹理图像和/深度图像来源于k个相机,一个相机属性信息字段用于存储一个相机的相机属性信息,因此,第i个轨道对应的视角信息中包括k 个相机属性信息字段用于存储k个相机的相机属性信息。相应的,一个相机属性字段包括上述1)-7),k个相机属性信息便包括k个上述1)-7)。
⑥第i个轨道对应的视角信息还包括图像信息字段,该图像信息字段用于存储第j个相机采集图像的图像信息。图像信息可以包括以下至少一种:深度图像降采样的倍数因子、深度图像左上顶点相对于平面帧原点的偏移量或纹理图像左上顶点相对于平面帧原点的偏移量。基于此,图像信息字段可以具体包括:1)深度图像降采样的倍数因子字段depth_downsample_factor,用于存储深度图像降采样的倍数因子;2)纹理图像左上顶点横轴偏移字段texture_vetex_x,用于存储纹理图像左上顶点相对于平面帧原点的偏移量横轴分量;3)纹理图像左上顶点纵轴偏移字段texture_vetex_y,用于存储纹理图像左上顶点相对于平面帧原点的偏移量纵轴分量;4)深度图像左上顶点横轴偏移字段depth_vetex_x,用于存储深度图像左上顶点相对于平面帧原点的横轴偏移量;5)深度图左上顶点纵轴偏移字段depth_vetex_y,用于存储深度图像左上顶点相对于平面帧原点的纵轴偏移量。
需要说明的是,一个图像信息字段用于存储一个相机采集图像的图像信息,第i个轨道内包括k个相机采集的深度图像和/或纹理图像,因此第i个轨道对应的自由视角信息中包括k个图像信息字段。
⑦第i个轨道对应的视角信息还包括自定义相机参数字段camera_parameter,该自定义相机参数字段用于存储第j个相机的第f个自定义相机参数,f为大于等于0且小于h的整数,h表示第j个相机的自定义相机参数的数量,第j个相机的自定义相机参数的数量可以存储在视角信息的自定义相机参数数量字段para_num中。
需要说明的是,由于第j个相机的自定义相机参数的数量为h个,一个自定义相机参数字段用于存储一个自定义相机参数,因此,第j个相机对应的自定相机参数字段的数量可以为h个。又因为,每个相机对应h个自定义相机参数,第i个轨道内包括k个相机,因此,第i个轨道对应的视角信息中包括k*h个自定义相机参数字段。
⑧第i个轨道对应的视角信息还包括自定义相机参数类型字段para_type,自定义相机参数类型字段para_type用于存储第j个相机的第f个自定义相机参数所属类型。需要说明的是,第j个相机的一个自定义相机参数所属类型存储在一个自定义相机参数类型字段中,由于第j个相机对应h个自定义相机参数,因此自定义相机参数类型字段的数量为h。与⑦同理的,第i个轨道对应的视角信息中包括k*h个自定义相机参数类型字段。
⑨第i个轨道对应的视角信息还包括自定义相机参数长度字段para_length,该自定义相机参数长度字段para_length用于存储第j个相机的第f个自定义相机参数的长度。由于第j个相机的一个自定义相机参数的长度存储在一个自定义相机参数长度字段中,因为第j个相机包括h个自定义相机参数,因此自定义相机参数长度字段的数量为h个。
本申请实施例中,S401中获取第i个轨道对应的自由视角信息数据盒是依据内容制作设备发送的信令描述文件实现的。具体地,在执行S401之前,获取沉浸媒体对应的信令描述文件,该信令描述文件包括沉浸媒体对应的自由视角相机描述子,自由视角相机描述子用于记录每个轨道内视频片段对应的相机属性信息,任一轨道内的视频片段是由任一轨道内的纹理图像和/或深度图像组成的;该与自由视角相机描述子被封装于所述媒体数据的媒体 呈现描述文件的自适应集层级中,或者所述信令描述文件被封装于所述媒体呈现描述文件的表示层级中。
本申请实施例中,自由视角相机描述子可以表示为AvsFreeViewCamInfo,其为SupplementalProperty元素,其@schemeIdUri属性为"urn:avs:ims:2018:av3l"。当自由视角相机描述子位于表示层级中时,可用于描述该表示representation层级对应的轨道内视频片段对应的相机属性信息;当自由视角相机描述子位于自适应层级adaptation set时,可以用于描述自适应层级中多个轨道内视频片段对应的相机属性信息。
其中,自由视角相机描述子中各个元素以及属性可如下表4所示:
表4
Figure PCTCN2022080257-appb-000005
进一步的,内容消费设备基于信令描述文件获取第i个轨道对应的自由视角信息数据盒。在一个实施例中,获取第i个轨道对应的自由视角信息数据盒,包括:基于自由视角相机描 述子中记录的每个轨道内图像对应的相机属性信息,从N个相机中选择与用户所在位置信息匹配的候选相机;向内容制作设备发送获取候选相机拍摄的图像的第一资源请求,第一资源请求用于指示内容制作设备根据M个轨道中每个轨道对应的自由视角信息数据盒中视角信息,从M个轨道中选择第i个轨道并返回第i个轨道对应的自由视角信息数据盒,第i个轨道中封装候选相机拍摄的图像;接收内容制作设备返回的第i个轨道对应的自由视角信息数据盒。在这种方式下,内容消费设备只需要获取所需的轨道对应的自由视角信息数据盒,不必获取所有轨道对应的自由视角信息数据盒,可节省传输资源。下面举例说明:
(1)假设内容制作设备生成自由视角视频并将自由视角视频封装为多个轨道,每个轨道中可以包括一个视角的纹理图像和深度图像,也就是说一个轨道内封装了来自一个相机的纹理图像和深度图像,一个轨道内的纹理图像和深度图像组成了一个视频片段,这样一来,可以理解为一个轨道内封装了来自处于一个视角的相机的视频片段。假设自由视角视频被封装到3个轨道,每个轨道内视频的纹理图像和深度图像组成一个视频片段,因此自由视角视频包括3个视频片段,分别表示为Representation1、Representation2以及Representation3;
(2)内容制作设备在信令生成环节,根据每个轨道对应的自由视角信息数据盒中的视角信息生成该自由视角视频对应的信令描述文件,信令描述文件中可携带自由视角相机描述子。假设自由视角视频自由视角相机描述子记录了3个视频片段对应的相机属性信息如下:
Representation1:{Cameral1:ID=1;Pos=(100,0,100);Focal=(10,20)};
Representation2:{Cameral2:ID=2;Pos=(100,100,100);Focal=(10,20)};
Representation1:{Cameral3:ID=3;Pos=(0,0,100);Focal=(10,20)};
(3)内容制作设备将信令描述文件发送至内容消费设备;
(4)内容消费设备根据信令描述文件和用户带宽,依据用户所在位置信息和信令描述文件中的相机属性信息,选取来自Cameral2和Cameral3的视频片段并向内容制作设备请求;
(5)内容制作设备将来自Cameral2和Cameral3的视频片段的轨道的自由视角信息数据盒发送至内容消费设备,内容消费设备根据获取到的轨道对应的自由视角信息数据盒初始化解码器,解码对应的视频片段并消费。
另一个实施例中,获取第i个轨道对应的自由视角信息数据盒,包括:基于信令描述文件和用户带宽向内容制作设备发送第二资源请求,第二资源请求用于指示内容制作设备返回M个轨道的M个自由视角信息数据盒,一个轨道对应一个自由视角信息数据盒;根据每个轨道对应的自由视角信息数据盒中的视角信息和用户所在位置信息,从M个自由视角信息数据盒中获取第i个轨道对应的自由视角信息数据盒。在这种方式下,内容消费设备虽然获取到所有轨道的自由视角信息数据盒,但是不对所有轨道中图像进行解码消费,只解码与用户当前位置匹配的第i个轨道中图像,节省了解码资源。下面举例来说:
(1)假设内容制作设备生成自由视角视频并将自由视角视频封装为多个轨道,每个轨道中可以包括一个视角的纹理图像和深度图像,也就是说每个轨道封装来自一个视角的相机的视频片段,假设自由视角视频被封装到3个轨道的视频片段分别表示为Representation1、Representation2以及Representation3;
(2)内容制作设备根据每个轨道对应的视角信息,生成该自由视角视频的信令描述文 件,如下所示:
Representation1:{Cameral1:ID=1;Pos=(100,0,100);Focal=(10,20)};
Representation2:{Cameral2:ID=2;Pos=(100,100,100);Focal=(10,20)};
Representation1:{Cameral3:ID=3;Pos=(0,0,100);Focal=(10,20)};
(3)内容制作设备将(2)中的信令描述文件发送至内容消费设备;
(4)内容消费设备根据信令描述文件和用户带宽向内容制作设备请求所有轨道对应的自由视角信息数据盒;假设自由视角视频被封装到3个轨道中,分别为Track1、Track2、以及Track3每个轨道中封装的视频片段对应的相机属性信息可以如下:
Track1:{Camera1:ID=1;Pos=(100,0,100);Focal=(10,20)};
Track2:{Camera2:ID=2;Pos=(100,100,100);Focal=(10,20)};
Track3:{Camera3:ID=3;Pos=(0,0,100);Focal=(10,20)};
(5)内容消费设备根据获取到的所有轨道对应的自由视角信息数据盒中的视角信息和用户当前观看的位置信息,选择track2和track3的轨道中封装的视频片段进行解码消费。
S402、根据自由视角信息数据盒中的视角信息对第i个轨道内封装的图像进行解码。
具体实现中,根据自由视角信息数据盒中的视角信息对第i个轨道内封装的图像进行解码显示,可以包括:根据第i个轨道对应的视角信息初始化解码器;再根据视角信息中与编码相关的指示信息对第i个轨道内封装的图像进行解码处理。
本申请实施例中将一沉浸媒体封装至M个轨道,该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中可以包括来自至少一个相机的图像,这M个轨道是属于同一个轨道组的,实现了一个沉浸视频封装到多个轨道的场景;另外,内容制作设备为每个轨道生成一个自由视角信息数据盒,通过第i个轨道对应的自由视角信息数据盒指示第i个轨道对应的视角信息,比如相机所在的具体视角位置,进而内容消费设备在根据第i个轨道对应的视角信息进行该轨道内图像的解码显示时,可以保证解码显示的视频与用户所在位置更加匹配,提高沉浸视频的呈现效果。
基于上述的沉浸媒体的数据处理方法,本申请实施例提供了另一种沉浸媒体的数据处理方法。参见图5,为本申请实施例提供的一种沉浸媒体的数据处理方法的流程示意图,图5所示的沉浸媒体的数据处理方法可由内容制作设备执行,具体可由内容制作设备的处理器执行。图5所示的沉浸媒体的数据处理方法可包括如下步骤:
S501、将沉浸媒体频封装到M个轨道中,该沉浸媒体是由处于不同视角的相机拍摄的图像组成的,一个轨道中封装了来自至少一个相机的图像,这M个轨道属于同一个轨道组。
S502、根据第i个轨道中w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒。
其中,w为大于或等于1的整数,第i个轨道内封装了w个图像,这w个图像可以是来自处于某一个或某几个视角的相机拍摄的深度图像和纹理图像中的至少一种。结合图4实施例中的代码段3,具体介绍根据w个图像在第i个轨道的封装过程生成第i个轨道对应的自由视角信息数据盒,具体可以包括:
①第i个轨道对应的自由视角信息数据盒中包括自由视角轨道组信息,该自由视角轨道组信息可以指示第i个轨道与其他封装了沉浸媒体的轨道属于同一个轨道组。该自由视角轨 道组信息包括第一相机标识字段camera_id和图像类型字段depth_texture_type,S502中根据第i个轨道内w个图像的封装过程生成所述第i个轨道对应的自由视角信息数据盒,包括:确定w个图像来源于k个相机,k为大于1的整数;将k个相机中的第j个相机的标识信息存储在第一相机标识字段camera_id中,j大于0且小于k;根据所述第j个相机采集的图像所属图像类型确定所述图像类型字段depth_texture_type,所述图像类型包括深度图像和纹理图像中任意一个或多个。应当理解的,一个相机的标识信息存储在一个第一相机标识字段camera_id,如果第i个轨道内包含的图像来源于k个相机,那么自由视角轨道组信息中可以包括k个第一相机标识字段camera_id。同理的,自由视角轨道组包括k个图像类型字段depth_texture_type。
具体地,根据第j个相机采集的图像所属图像类型确定图像类型字段,包括:当第j个相机采集的图像所属图像类型为纹理图像,则将图像类型字段设置为第二数值;当第j个相机采集的图像所属图像类型为深度图像,则将图像类型字段设置为第三数值;当第j个相机采集的图像所属图像类型为纹理图像和深度图像,则将图像类型字段设置为第四数值。其中,第二数值可以为1,第三数值可以为2,第四数值可以3,图像类型字段可参见图4实施例中表2所示。
②第i个轨道对应的视角信息包括视频拼接布局指示字段stitching_layout,S502中根据第i个轨道内w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:如果第i个轨道内包含的纹理图像和深度图像是拼接编码的,则将视频布局指示字段stitching_layout设置为第一数值;如果第i个轨道内包含的纹理图像和深度图像是分别编码的,则将视频布局指示字段stitching_layout设置为第二数值。其中,第一数值可以为0,第二数值可以为1,视频拼接布局指示字段可以参见图4实施例中表3所示。
③第i个轨道对应的视角信息还包括相机模型字段camera_model;S502中根据第i个轨道内w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:当采集w个图像的相机所属的相机模型为第一模型时,将相机模型字段camera_model设置为第一数值;当采集w个图像的相机所属的相机模型为第二模型时,将相机模型字段camera_model设置为第二数值。同上述,第一数值可以为0,第二数值可以为1,相机模型字段可以参见图4实施例中表4所示。
④第i个轨道对应的视角信息还包括纹理图像的保护带宽字段texture_padding_size和深度图像的保护带宽字段depth_padding_size。S502中根据第i个轨道内w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:获取第i个轨道内对纹理图像编码时采用的纹理图像的保护带宽度,并将纹理图像的保户带宽度存储在纹理图像的保护带宽度字段texture_padding_size;获取第i个轨道内对深度图像编码时采用的深度图像的保护带宽度,并将深度图像的保护带宽度存储在深度图像的保护带宽度字段depth_padding_size。
⑤第i个轨道对应的视角信息还包括第二相机标识字段camera_id和相机属性信息字段,S502中根据第i个轨道内w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:将第j个相机的标识信息存储在第二相机标识字段camera_id,j大于或等于0且小于k,k表示w个图像来源相机的数量;获取第j个相机的相机属性信息,并将获取到的相机属性信息存储在相机属性信息字段。其中,所述第j个相机的相机属性信息包括以下任意一种或多 种:第j个相机位置的横轴分量取值、纵轴分量取值以及竖轴分量取值,第j个相机焦距的横轴分量取值和纵轴分量取值,以及第j个相机采集图像的分辨率宽度与高度。
基于此,相机属性信息字段可具体包括:1)相机位置的横轴分量字段camera_pos_x,用于存储第j个相机位置的横轴分量取值(也称为x分量取值);2)相机位置的纵轴分量字段camera_pos_y,用于存储第j个相机位置的纵轴分量取值(也称为y分量取值);3)相机位置的竖轴分量字段camera_pos_z,用于存储第j个相机位置的竖轴分量取值(也称为z分量取值);4)相机焦距的横轴分量字段focal_length_x,用于存储第j个相机焦距的横轴分量取值(也称为x分量取值);5)相机焦距的纵轴分量字段focal_length_y,用于存储第j个相机焦距的纵轴分量取值(也称为y分量取值);6)相机采集图像的分辨率宽度字段camera_resolution_x,用于存储第j个相机采集图像的分辨率宽度;7)相机采集图像的分辨率宽度字段camera_resolution_y,用于存储第j个相机采集图像的分辨率高度。
需要说明的是,由于第i个轨道内的纹理图像和/深度图像来源于k个相机,一个相机属性信息字段用于存储一个相机的相机属性信息,因此,第i个轨道对应的视角信息中包括k个相机属性信息字段用于存储k个相机的相机属性信息。相应的,一个相机属性字段包括上述1)-7),k个相机属性信息便包括k个上述1)-7)。
⑥第i个轨道对应的视角信息还包括图像信息字段,S502中根据w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:获取所述第j个相机采集图像的图像信息,并将获取到的图像信息存储在图像信息字段;其中,所述图像信息包括如下一个或多个:深度图像降采样的倍数因子、深度图像左上顶点相对于平面帧原点的偏移量、纹理图像左上顶点相对于平面帧原点的偏移量。
基于此,图像信息字段可以具体包括:1)深度图像降采样的倍数因子字段depth_downsample_factor,用于存储深度图像降采样的倍数因子;2)纹理图像左上顶点横轴偏移字段texture_vetex_x,用于存储纹理图像左上顶点相对于平面帧原点的偏移量横轴分量;3)纹理图像左上顶点纵轴偏移字段texture_vetex_y,用于存储纹理图像左上顶点相对于平面帧原点的偏移量纵轴分量;4)深度图像左上顶点横轴偏移字段depth_vetex_x,用于存储深度图像左上顶点相对于平面帧原点的横轴偏移量;5)深度图左上顶点纵轴偏移字段depth_vetex_y,用于存储深度图像左上顶点相对于平面帧原点的纵轴偏移量。
需要说明的是,一个图像信息字段用于存储一个相机采集图像的图像信息,第i个轨道内包括k个相机采集的深度图像和/或纹理图像,因此第i个轨道对应的自由视角信息中包括k个图像信息字段。
⑦第i个轨道对应的视角信息还包括自定义相机参数字段、自定义相机参数类型字段以及自定义相机参数长度字段;根据第i个轨道内w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:获取第f个自定义相机参数,并将第f个自定义相机参数存储在自定义相机参数字段,f为大于等于0且小于h的整数,h表示第j个相机的自定义相机参数的数量;确定第f个自定义相机参数所属的参数类型以及第f个自定义相机参数的长度,将第f个自定义相机参数所属的参数类型存储在自定义相机参数类型字段,以及将第f个自定义相机参数的长度存储在自定义相机参数长度字段。
需要说明的是,由于第j个相机的自定义相机参数的数量为h个,一个自定义相机参数字段用于存储一个自定义相机参数,因此,第j个相机对应的自定相机参数字段的数量可以为h个。又因为,每个相机对应h个自定义相机参数,第i个轨道内包括k个相机,因此,第i个轨道对应的视角信息中包括k*h个自定义相机参数字段。需要说明的是,由于第j个相机的自定义相机参数的数量为h个,一个自定义相机参数字段用于存储一个自定义相机参数,因此,第j个相机对应的自定相机参数字段的数量可以为h个。又因为,每个相机对应h个自定义相机参数,第i个轨道内包括k个相机,因此,第i个轨道对应的视角信息中包括k*h个自定义相机参数字段。由于第j个相机的一个自定义相机参数的长度存储在一个自定义相机参数长度字段中,因为第j个相机包括h个自定义相机参数,因此自定义相机参数长度字段的数量为h个。
另外,内容制作设备还可以根据M个轨道对应的自由视角信息数据盒中的视角信息生成信令描述文件,信令描述文件包括沉浸媒体对应的自由视角相机描述子,自由视角相机描述子用于记录每个轨道内图像对应的相机属性信息;自由视角相机描述子被封装于沉浸媒体的媒体呈现描述文件的自适应集层级中,或者信令描述文件被封装于媒体呈现描述文件的表示层级中。
进一步的,内容制作设备可以将信令描述文件发送至内容消费设备,以使内容消费设备根据信令描述文件获取第i个轨道对应的自由视角信息数据盒。
作为一种可选的实施方式,内容制作设备将信令描述文件发送至内容消费设备,以指示内容消费设备基于自由视角相机描述子中记录的每个轨道内图像对应的相机属性信息,从N个相机中选择与用户所在位置匹配的候选相机,以及发送获取来源于候选相机的图像的第一资源请求;响应于第一资源请求,根据M个轨道中每个轨道对应的自由视角信息数据盒中的视角信息从M个轨道中选择第i个轨道并将i个轨道对应的自由视角信息数据盒发送至内容消费设备,第i个轨道内包括来源于候选相机的图像。
作为另一种可选的实施方式,内容制作设备将信令描述文件发送至内容消费设备,以指示内容消费设备根据信令描述文件和用户带宽发送第二资源请求;响应于第二资源请求,将M个轨道对应的M个自由视角信息数据盒发送至内容消费设备,以指示内容消费设备根据每个轨道对应的自由视角信息数据盒中的视角信息和用户所在位置信息,从M个自由视角信息数据盒中获取第i个轨道对应的自由视角信息数据盒。
本申请实施例中将一个沉浸媒体封装至M个轨道,该沉浸视频是由处于不同视角的相机采集的图像组成的,这M个轨道是属于同一个轨道组的,一个轨道中封装来自至少一个相机的图像,实现了一个沉浸视频封装到多个轨道的场景;另外,内容制作设备为根据每个轨道中图像的封装过程为每个轨道生成一个自由视角信息数据盒,通过第i个轨道对应的自由视角信息数据盒指示第i个轨道对应的视角信息,比如相机所在的具体视角位置,进而内容消费设备在根据第i个轨道对应的视角信息进行该轨道内图像的解码显示时,可以保证解码显示的视频与用户所在位置更加匹配,提高沉浸视频的呈现效果。
基于上述的沉浸媒体的数据处理方法实施例,本申请实施例提供了一种沉浸媒体的数据处理装置,该沉浸媒体的数据处理装置可以是运行于内容消费设备中的一个计算机程序 (包括程序代码),例如该沉浸媒体的数据处理装置可以是内容消费设备中的一个应用软件。参见图6,为本申请实施例提供的一种沉浸媒体的数据处理装置的结构示意图。图6所示的数据处理装置可运行如下单元:
获取单元601,用于获取所述沉浸媒体的第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒包括所述第i个轨道对应的视角信息,i为大于或等于1且小于或等于M的整数;该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,所述沉浸媒体被封装到M个轨道中,一个轨道中封装来自至少一个相机的图像,所述M个轨道属于同一个轨道组,N和M均为大于1的整数;
处理单元602,用于根据所述自由视角信息数据盒中的视角信息对所述第i个轨道内封装的图像进行解码。
在一个实施例中,所述第i个轨道内封装了来自所述N个相机中k个相机的图像,k为大于0的整数;所述第i个轨道内还封装了所述第i个轨道对应的自由视角轨道组信息,所述自由视角轨道组信息用于指示所述第i个轨道与封装所述沉浸媒体的其他轨道属于同一个轨道组;所述自由视角轨道组信息包括第一相机标识字段和图像类型字段;
所述第一相机标识字段用于存储所述k个相机中第j个相机的标识信息,所述图像类型字段用于指示所述第j个相机采集的图像所属的图像类型,所述图像类型包括纹理图像或深度图像中的至少一个。
在一个实施例中,若所述第i个轨道内包括纹理图像和深度图像,所述第i个轨道对应的视角信息包括视频拼接布局指示字段;当所述视频拼接布局指示字段为第一数值时,指示所述第i个轨道内包括的纹理图像和深度图像是拼接编码的;当所述视频拼接布局指示字段为第二数值时,指示所述第i个轨道内包括的纹理图像和深度图像是分别编码的。
在一个实施例中,所述第i个轨道对应的视角信息还包括相机模型字段;当所述相机模型字段为第三数值时,指示所述j个相机所属相机模型为第一模型;当所述相机模型字段为第四数值时,指示所述j个相机所属相机模型为第二模型。
在一个实施例中,所述第i个轨道对应的视角信息还包括纹理图像的保护带宽度字段和深度图像的保护带宽度字段,所述纹理图像的保护带宽度字段用于存储所述第i个轨道内纹理图像进行编码时采用的保护带宽度,所述深度图像的保护带宽度字段用于存储所述第i个轨道内深度图像进行编码时采用的保护带宽度。
在一个实施例中,所述第i个轨道对应的视角信息还包括第二相机标识字段和相机属性信息字段;所述第二相机标识字段用于存储第j个相机的标识信息;
所述相机属性信息字段用于存储所述第j个相机的相机属性信息,所述第j个相机的相机属性信息包括以下至少一种:所述第j个相机位置的横轴分量取值、纵轴分量取值以及竖轴分量取值,所述第j个相机焦距的横轴分量取值和纵轴分量取值,以及第j个相机采集图像的分辨率宽度与高度。
在一个实施例中,所述第i个轨道对应的视角信息还包括图像信息字段,所述图像信息字段用于存储所述第j个相机采集图像的图像信息,所述图像信息包括以下任意一个:深度图像降采样的倍数因子、深度图像左上顶点相对于平面帧原点的偏移量或纹理图像左上顶 点相对于平面帧原点的偏移量。
在一个实施例中,所述第i个轨道对应的视角信息还包括自定义相机参数字段、自定义相机参数类型字段以及自定义相机参数长度字段;所述自定义相机参数字段用于存储所述第j个相机的第f个自定义相机参数,f为大于等于0且小于h的整数,h表示所述第j个相机的自定义相机参数的数量;所述自定义相机参数类型字段用于存储所述第f个自定义相机参数所属的参数类型;所述自定义相机参数长度字段用于存储所述第f个自定义相机参数的长度。
在一个实施例中,所述获取单元601还用于:
获取所述沉浸媒体对应的信令描述文件,所述信令描述文件包括所述沉浸媒体对应的自由视角相机描述子,所述自由视角相机描述子用于记录每个轨道内视频片段对应的相机属性信息,所述轨道内视频片段是由所述轨道内包括的纹理图像和深度图像组成的;所述自由视角相机描述子被封装于所述沉浸媒体的媒体呈现描述文件的自适应集层级中,或者所述信令描述文件被封装于所述媒体呈现描述文件的表示层级中。
在一个实施例中,所述获取单元601在获取所述沉浸媒体的第i个轨道对应的自由视角信息数据盒时,执行如下操作:
基于所述自由视角相机描述子中记录的每个轨道内视频片段对应的相机属性信息,从所述N个相机中选择与用户所在位置信息匹配的候选相机;
向内容制作设备发送获取来源于所述候选相机的分块视频的第一资源请求,所述第一资源请求用于指示所述内容制作设备根据所述M个轨道中每个轨道对应的自由视角信息数据盒中的视角信息,从所述M个轨道中选择第i个轨道并返回所述第i个轨道对应的自由视角信息数据盒,所述第i个轨道中封装的纹理图像和深度图像中的至少一种来自所述候选相机;接收所述内容制作设备返回的所述第i个轨道对应的自由视角信息数据盒。
在一个实施例中,所述获取单元601在获取获取所述沉浸媒体的第i个轨道对应的自由视角信息数据盒时,执行如下操作:
基于所述信令描述文件和用户带宽向内容制作设备发送第二资源请求,所述第二资源请求用于指示所述内容制作设备返回所述M个轨道的M个自由视角信息数据盒,一个轨道对应一个自由视角信息数据盒;
根据每个轨道对应的自由视角信息数据盒中的视角信息和用户所在位置信息,从所述M个自由视角信息数据盒中获取第i个轨道对应的自由视角信息数据盒。
根据本申请的一个实施例,图4所示的沉浸媒体的数据处理方法所涉及各个步骤可以是由图6所示的沉浸媒体的数据处理装置中的各个单元来执行的。例如,图4所述的S401可由图6中所述的数据处理装置中的获取单元601来执行,S402可由图6所示的数据处理装置中的处理单元602来执行。
根据本申请的另一个实施例,图6所示的沉浸媒体的数据处理装置中的各个单元可以分别或全部合并为一个或若干个另外的单元来构成,或者其中的某个(些)单元还可以再拆分为功能上更小的多个单元来构成,这可以实现同样的操作,而不影响本申请的实施例的技术效果的实现。上述单元是基于逻辑功能划分的,在实际应用中,一个单元的功能也可以由多个单元来实现,或者多个单元的功能由一个单元实现。在本申请的其它实施例中, 基于沉浸媒体的数据处理装置也可以包括其它单元,在实际应用中,这些功能也可以由其它单元协助实现,并且可以由多个单元协作实现。
根据本申请的另一个实施例,可以通过在包括中央处理单元(CPU)、随机存取存储介质(RAM)、只读存储介质(ROM)等处理元件和存储元件的例如计算机的通用计算设备上运行能够执行如图4的相应方法所涉及的各步骤的计算机程序(包括程序代码),来构造如图6中所示的沉浸媒体的数据处理装置,以及来实现本申请实施例沉浸媒体的数据处理方法。所述计算机程序可以记载于例如计算机可读存储介质上,并通过计算机可读存储介质装载于上述计算设备中,并在其中运行。
本申请实施例中将一沉浸媒体封装至M个轨道,该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中可以包括来自至少一个相机的图像,这M个轨道是属于同一个轨道组的,实现了一个沉浸视频封装到多个轨道的场景;另外,内容制作设备为每个轨道生成一个自由视角信息数据盒,通过第i个轨道对应的自由视角信息数据盒指示第i个轨道对应的视角信息,比如相机所在的具体视角位置,进而内容消费设备在根据第i个轨道对应的视角信息进行该轨道内图像的解码显示时,可以保证解码显示的视频与用户所在位置更加匹配,提高沉浸视频的呈现效果。
基于上述的沉浸媒体的数据处理方法以及数据处理装置实施例,本申请实施例提供了另一种沉浸媒体的数据处理装置,该沉浸媒体的数据处理装置可以是运行于内容制作设备中的一个计算机程序(包括程序代码),例如该沉浸媒体的数据处理装置可以是内容制作设备中的一个应用软件。参见图7,为本申请实施例提供的另一种沉浸媒体的数据处理装置的结构示意图。图7所示的数据处理装置可运行如下单元:
封装单元701,用于将沉浸媒体封装到M个轨道中,所述沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中封装来自至少一个相机的图像,所述M个轨道属于同一个轨道组,N和M均为大于或等于1的整数;
生成单元702,用于根据第i个轨道中w个图像的封装过程,生成第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒包括所述第i个轨道对应的视角信息;1≤i≤M,w≥1。
在一个实施例中,所述第i个轨道对应的自由视角信息数据盒包括自由视角轨道组信息,所述自由视角轨道组信息包括第一相机标识字段和图像类型字段,所述封装单元701在根据第i个轨道内w个图像的封装过程生成所述第i个轨道对应的自由视角信息数据盒时,执行如下步骤:
确定所述w个图像来源于k个相机,k为大于1的整数;将所述k个相机中的第j个相机的标识信息存储在所述第一相机标识字段中;根据所述第j个相机采集的图像所属图像类型确定所述图像类型字段,所述图像类型包括深度图像或纹理图像中的至少一个。
在一个实施例中,所述封装单元701在根据所述第j个相机采集的图像所属图像类型确定所述图像类型字段时,执行如下步骤:
当所述第j个相机采集的图像所属图像类型为纹理图像,则将所述图像类型字段设置为第二数值;当所述第j个相机采集的图像所属图像类型为深度图像,则将所述图像类型字段 设置为第三数值;当所述第j个相机采集的图像所属图像类型为纹理图像和深度图像,则将所述图像类型字段设置为第四数值。
在一个实施例中,所述第i个轨道对应的视角信息包括视频拼接布局指示字段,所述w个图像包括纹理图像和深度图像,所述封装单元701在根据第i个轨道内w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒时,执行如下步骤:
如果所述第i个轨道内包含的纹理图像和深度图像是拼接编码的,则将所述视频布局指示字段设置为第一数值;如果所述i个轨道内包含的的纹理图像和深度图像是分别编码的,则将所述视频布局指示字段设置为第二数值。
在一个实施例中,所述第i个轨道对应的视角信息还包括相机模型字段;所述封装单元701在根据第i个轨道内的w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒时,执行如下步骤:
当采集所述w个图像的相机所属的相机模型为第一模型时,将所述相机模型字段设置为第一数值;当采集所述w个图像的相机所属的相机模型为第二模型时,将所述相机模型字段设置为第二数值。
在一个实施例中,所述第i个轨道对应的视角信息还包括纹理图像的保护带宽字段和深度图像的保护带宽字段,所述封装单元701在根据第i个轨道内的w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒时,执行如下步骤:
获取所述第i个轨道内在对纹理图像进行编码时采用的纹理图像的保护带宽度,并将所述纹理图像的保户带宽度存储在所述纹理图像的保护带宽度字段;
获取所述第i个轨道内在对深度图像进行编码时采用的深度图像的保护带宽度,并将所述深度图像的保护带宽度存储在所述深度图像的保护带宽度字段。
在一个实施例中,所述第i个轨道对应的视角信息还包括第二相机标识字段和相机属性信息字段,所述封装单元701在根据第i个轨道内的w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒时,执行如下步骤:
将k个相机中第j个相机的标识信息存储在所述第二相机标识字段,k表示所述w个图像来源相机的数量;
获取所述第j个相机的相机属性信息,并将获取到的相机属性信息存储在所述相机属性信息字段;
所述第j个相机的相机属性信息包括以下任意一种或多种:所述第j个相机位置的横轴分量取值、纵轴分量取值以及竖轴分量取值,所述第j个相机焦距的横轴分量取值和纵轴分量取值,以及所述第j个相机采集图像的分辨率宽度与高度。
在一个实施例中,所述第i个轨道对应的视角信息还包括图像信息字段,所述封装单元701在根据第i个轨道内的w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒时,执行如下步骤:
获取所述第j个相机采集的图像的图像信息,并将获取到的图像信息存储在所述图像信息字段;其中,所述图像信息包括如下一个或多个:深度图像降采样的倍数因子、深度图像左上顶点相对于平面帧原点的偏移量、纹理图像左上顶点相对于平面帧原点的偏移量。
在一个实施例中,所述第i个轨道对应的视角信息还包括自定义相机参数字段、自定义相机参数类型字段以及自定义相机参数长度字段;所述封装单元701在根据第i个轨道内的w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒时,执行如下步骤:
获取第f个自定义相机参数,并将所述第f个自定义相机参数存储在所述自定义相机参数字段,f为大于等于0且小于h的整数,h表示所述第j个相机的自定义相机参数的数量;
确定所述第f个自定义相机参数所属的参数类型以及所述第f个自定义相机参数的长度,将所述第f个自定义相机参数所属的参数类型存储在所述自定义相机参数类型字段,以及将所述第f个自定义相机参数的长度存储在所述自定义相机参数长度字段。
在一个实施例中,所述生成单元702还用于:生成所述沉浸媒体对应的信令描述文件,所述信令描述文件包括所述沉浸媒体对应的自由视角相机描述子,所述自由视角相机描述子用于记录每个轨道内视频片段对应的相机属性信息,任一轨道内视频片段是由所述任一轨道内封装的图像组成的;所述自由视角相机描述子被封装于所述沉浸媒体的媒体呈现描述文件的自适应集层级中,或者所述信令描述文件被封装于所述媒体呈现描述文件的表示层级中。
在一个实施例中,所述沉浸媒体的数据处理装置还包括发送单元703,用于将所述信令描述文件发送至内容消费设备,以指示所述内容消费设备基于所述自由视角相机描述子中记录的每个轨道内视频片段对应的相机属性信息,从所述多个相机中选择与用户所在位置匹配的候选相机,以及发送获取来源于所述候选相机的分块视频的第一资源请求;响应于所述第一资源请求,根据所述M个轨道中每个轨道对应的自由视角信息数据盒中的视角信息从所述M个轨道中选择第i个轨道并将所述i个轨道对应的自由视角信息数据盒发送至所述内容消费设备。
在一个实施例中,所述发送单元703还用于:将所述信令描述文件发送至内容消费设备,以指示所述内容消费设备根据所述信令描述文件和用户带宽发送第二资源请求;响应于所述第二资源请求,将所述M个轨道对应的M个自由视角信息数据盒发送至所述内容消费设备,以指示所述内容消费设备根据每个轨道对应的自由视角信息数据盒中的视角信息和用户所在位置信息,从所述M个自由视角信息数据盒中获取第i个轨道对应的自由视角信息数据盒。
根据本申请的一个实施例,图5所示的沉浸媒体的数据处理方法所涉及各个步骤可以是由图7所示的沉浸媒体的数据处理装置中的各个单元来执行的。例如,图5所述的S501可由图7中所述的数据处理装置中的封装单元702来执行,S502可由图7所述的数据处理装置中的生成单元702来执行。
根据本申请的另一个实施例,图7所示的沉浸媒体的数据处理装置中的各个单元可以分别或全部合并为一个或若干个另外的单元来构成,或者其中的某个(些)单元还可以再拆分为功能上更小的多个单元来构成,这可以实现同样的操作,而不影响本申请的实施例的技术效果的实现。上述单元是基于逻辑功能划分的,在实际应用中,一个单元的功能也可以由多个单元来实现,或者多个单元的功能由一个单元实现。在本申请的其它实施例中,基于沉浸媒体的数据处理装置也可以包括其它单元,在实际应用中,这些功能也可以由其 它单元协助实现,并且可以由多个单元协作实现。
根据本申请的另一个实施例,可以通过在包括中央处理单元(CPU)、随机存取存储介质(RAM)、只读存储介质(ROM)等处理元件和存储元件的例如计算机的通用计算设备上运行能够执行如图5的相应方法所涉及的各步骤的计算机程序(包括程序代码),来构造如图7中所示的沉浸媒体的数据处理装置,以及来实现本申请实施例沉浸媒体的数据处理方法。所述计算机程序可以记载于例如计算机可读存储介质上,并通过计算机可读存储介质装载于上述计算设备中,并在其中运行。
本申请实施例中将一沉浸媒体封装至M个轨道,该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中可以包括来自至少一个相机的图像,这M个轨道是属于同一个轨道组的,实现了一个沉浸视频封装到多个轨道的场景;另外,内容制作设备为每个轨道生成一个自由视角信息数据盒,通过第i个轨道对应的自由视角信息数据盒指示第i个轨道对应的视角信息,比如相机所在的具体视角位置,进而内容消费设备在根据第i个轨道对应的视角信息进行该轨道内图像的解码显示时,可以保证解码显示的视频与用户所在位置更加匹配,提高沉浸视频的呈现效果。
参见图8,为本申请实施例提供的一种内容消费设备的结构示意图。图8所示的内容消费设备可以指沉浸媒体的使用者所使用的计算机设备,该计算机设备可以是终端。图8所示的内容消费设备可以包括接收器801、处理器802、存储器803以及显示/播放装置804。其中:
接收器801用于实现解码与其他设备的传输交互,具体用于实现内容制作设备与内容消费设备之间关于沉浸媒体的传输。即内容消费设备通过接收器901接收沉内容制作设备传输沉浸媒体的相关媒体资源。
处理器802或称CPU(Central Processing Unit,中央处理器))是内容制作设备的处理核心,该处理器802适于实现一条或多条计算机程序,具体适于加载并执行一条或多条计算机程序从而实现图4所示的沉浸媒体的数据处理方法的流程。
存储器803是内容消费设备中的记忆设备,用于存储计算机程序和媒体资源。可以理解的,此处的存储器803既可以包括内容消费设备中的内置存储介质,当然也可以包括内容消费设备所支持的扩展存储介质。需要说明的是,存储器803可以是高速RAM存储器,也可以是非不稳定的存储器(non-volatile memory),例如至少一个磁盘存储器;可选的还可以是至少一个位于远离前述处理器的存储器。存储器803提供存储空间,该存储空间用于存储内容消费设备的操作系统。并且,在该存储空间中还用于存储计算机程序,该计算机程序适于被处理器调用并执行,以用来执行沉浸媒体的数据处理方法的各步骤。另外,存储器803还可用于存储经处理器处理后形成的沉浸媒体的三维图像、三维图像对应的音频内容及该三维图像和音频内容渲染所需的信息等。
显示/播放装置804用于输出渲染得到的声音和三维图像。
在一个实施例中,处理器802可包括解析器821、解码器822、转换器823以及渲染器824;其中:
解析器821用于对来自内容制作设备的渲染媒体的封装文件进行解封装,具体是按照沉浸媒体的文件格式要对媒体文件资源进行解封装,得到音频码流和视频码流;并将该音频 码流和视频码流提高给解码器822;
解码器822对音频码流进行音频解码,得到音频内容并提供给渲染器824进行音频渲染。另外,解码器822对视频码流进行解码得到2D图像。根据媒体呈现描述信息提供的元数据,如果该元数据指示沉浸媒体执行过区域封装过程,该2D图像是指封装图像;如果该元数据指示沉浸媒体未执行过区域封装过程,则该平面图像是指投影图像。
转换器823用于将2D图像转换为3D图像。如果沉浸媒体执行过区域封装过程,转换器923还会先将封装图像进行区域解封装得到投影图像。再对投影图像进行重建处理得到3D图像。如果渲染媒体未执行过区域封装过程,转换器923会直接将投影图像重建得到3D图像。
渲染器824用于对沉浸媒体的音频内容和3D图像进行渲染。具体根据媒体呈现描述信息中与渲染、视窗相关的元数据对音频内容及3D图像进行渲染,渲染完成交由显示/播放装置进行输出。
在一个实施例中,处理器802通过调用存储器中的一条或多条计算机程序执行图4所示的沉浸媒体的数据处理方法的各个步骤。具体地,存储器存储一条或多条计算机程序,该一条或多条计算机程序适于由处理器802加载并执行如下步骤:
获取第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒子,包括所述第i个轨道对应的视角信息,i为大于等于1且小于等于M的整数;该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成,所述沉浸媒体被封装到M个轨道,一个轨道中封装来自至少一个相机的图像,N和M均为大于1的整数。
根据所述自由视角信息数据盒中的视角信息对所述第i个轨道内封装的图像进行解码显示。
本申请实施例中将一沉浸媒体封装至M个轨道,该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中可以包括来自至少一个相机的图像,这M个轨道是属于同一个轨道组的,实现了一个沉浸视频封装到多个轨道的场景;另外,内容制作设备为每个轨道生成一个自由视角信息数据盒,通过第i个轨道对应的自由视角信息数据盒指示第i个轨道对应的视角信息,比如相机所在的具体视角位置,进而内容消费设备在根据第i个轨道对应的视角信息进行该轨道内图像的解码显示时,可以保证解码显示的视频与用户所在位置更加匹配,提高沉浸视频的呈现效果。
参见图9,为本申请实施例提供的一种内容制作设备的结构示意图;该内容制作设备可以是指沉浸媒体的提供者所使用的计算机设备,该计算机设备可以是终端或服务器。如图9所示,该内容制作设备可以包括捕获设备901、处理器902、存储器903以及发射器904。其中:
捕获设备901用于采集现实世界的声音-视觉场景获得沉浸媒体的原始数据(包括在时间和空间上保持同步的音频内容和视频内容)。该捕获设备801可以包括但不限于:音频设备、摄像设备及传感设备。其中,音频设备可以包括音频传感器、麦克风等。摄像设备可以包括普通摄像头、立体摄像头、光场摄像头等。传感设备可以包括激光设备、雷达设备等。
处理器902(或称CPU(Central Processing Unit,中央处理器))是内容制作设备的处理核心,该处理器902适于实现一条或多条计算机程序,具体适于加载并执行一条或多条计算机程序从而实现图4所示的沉浸媒体的数据处理方法的流程。
存储器903是内容制作设备中的记忆设备,用于存放程序和媒体资源。可以理解的是,此处的存储器903既可以包括内容制作设备中的内置存储介质,当然也可以包括内容制作设备所支持的扩展存储介质。需要说明的是,存储器可以是高速RAM存储器,也可以是非不稳定的存储器(non-volatile memory),例如至少一个磁盘存储器;可选的还可以是至少一个位于远离前述处理器的存储器。存储器提供存储空间,该存储空间用于存储内容制作设备的操作系统。并且,在该存储空间中还用于存储计算机程序,该计算机程序包括程序指令,且该程序指令适于被处理器调用并执行,以用来执行沉浸媒体的数据处理方法的各步骤。另外,存储器903还可用于存储经处理器处理后形成的沉浸媒体文件,该沉浸媒体文件包括媒体文件资源和媒体呈现描述信息。
发射器904用于实现内容制作设备与其他设备的传输交互,具体用于实现内容制作设备与内容消费设备之间关于进行沉浸媒体的传输。即内容制作设备通过发射器904来向内容消费设备传输沉浸媒体的相关媒体资源。
再参见图9可知,处理器902可以包括转换器921、编码器922和封装器923。其中:
转换器921用于对捕获到的视频内容进行一系列转换处理,使视频内容成为适合被执行沉浸媒体的视频编码的内容。转换处理可包括:拼接和投影,可选地,转换处理还包括区域封装。转换器921可以将捕获到的3D视频内容转换为2D图像,并提供给编码器进行视频编码。
编码器922用于对捕获到的音频内容进行音频编码形成沉浸媒体的音频码流。还用于对转换器921转换得到的2D图像进行视频编码,得到视频码流。
封装器923用于将音频码流和视频码流按照沉浸媒体的文件格式(如ISOBMFF)封装在文件容器中形成沉浸媒体的媒体文件资源,该媒体文件资源可以是媒体文件或媒体片段形成沉浸媒体的媒体文件;并按照沉浸媒体的文件格式要求采用媒体呈现描述信息记录该沉浸媒体的媒体文件资源的元数据。封装器处理得到的沉浸媒体的封装文件会保存在存储器中,并按需提供给内容消费设备进行沉浸媒体的呈现。
在一个实施例中,处理器902通过调用存储器中的一条或多条指令来执行图5所示的沉浸媒体的数据处理方法的各个步骤。具体地,存储器803存储有一条或多条计算机程序,该计算机程序适于由处理器902加载并执行但不限于如下步骤:
将所述沉浸媒体封装到M个轨道中,所述沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中封装来自至少一个相机的图像,所述M个轨道属于同一个轨道组,N和M均为大于或等于1的整数;
根据第i个轨道中w个图像的封装过程,生成第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒包括所述第i个轨道对应的视角信息;1≤i≤M,w≥1。
本申请实施例中将一沉浸媒体封装至M个轨道,该沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中可以包括来自至少一个相机的图像,这M个轨道是属于 同一个轨道组的,实现了一个沉浸视频封装到多个轨道的场景;另外,内容制作设备为每个轨道生成一个自由视角信息数据盒,通过第i个轨道对应的自由视角信息数据盒指示第i个轨道对应的视角信息,比如相机所在的具体视角位置,进而内容消费设备在根据第i个轨道对应的视角信息进行该轨道内图像的解码显示时,可以保证解码显示的视频与用户所在位置更加匹配,提高沉浸视频的呈现效果。
另外,本申请实施例还提供了一种存储介质,所述存储介质用于存储计算机程序,所述计算机程序用于执行上述实施例提供的方法。
本申请实施例还提供了一种包括指令的计算机程序产品,当其在计算机上运行时,使得计算机执行上述实施例提供的方法。
以上所揭露的仅为本申请较佳实施例而已,当然不能以此来限定本申请之权利范围,因此依本申请权利要求所作的等同变化,仍属本申请所涵盖的范围。

Claims (28)

  1. 一种沉浸媒体的数据处理方法,所述方法由内容消费设备执行,所述沉浸媒体是由处于不同视角的N个相机拍摄的图像组成,所述沉浸媒体被封装到M个轨道,一个轨道中封装来自至少一个相机的图像,N和M均为大于1的整数;所述方法包括:
    获取所述沉浸媒体的第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒包括所述第i个轨道对应的视角信息,i为大于或等于1且小于或等于M的整数;
    根据所述自由视角信息数据盒中的视角信息对所述第i个轨道内封装的图像进行解码。
  2. 如权利要求1所述的方法,所述第i个轨道内封装了来自所述N个相机中k个相机的图像,k为大于0的整数;
    所述第i个轨道内还封装了所述第i个轨道对应的自由视角轨道组信息,所述自由视角轨道组信息用于指示所述第i个轨道与封装所述沉浸媒体的其他轨道属于同一个轨道组;所述自由视角轨道组信息包括第一相机标识字段和图像类型字段;
    所述第一相机标识字段用于存储所述k个相机中第j个相机的标识信息,所述图像类型字段用于指示所述第j个相机采集的图像所属的图像类型,所述图像类型包括纹理图像或深度图像中的至少一个。
  3. 如权利要求2所述的方法,若所述第i个轨道内包括纹理图像和深度图像,所述第i个轨道对应的视角信息包括视频拼接布局指示字段;
    当所述视频拼接布局指示字段为第一数值时,指示所述第i个轨道内包括的纹理图像和深度图像是拼接编码的;
    当所述视频拼接布局指示字段为第二数值时,指示所述第i个轨道内包括的纹理图像和深度图像是分别编码的。
  4. 如权利要求2所述的方法,所述第i个轨道对应的视角信息还包括相机模型字段;当所述相机模型字段为第三数值时,指示所述j个相机所属相机模型为第一模型;当所述相机模型字段为第四数值时,指示所述j个相机所属相机模型为第二模型,所述第二模型不同于所述第一模型。
  5. 如权利要求2所述的方法,所述第i个轨道对应的视角信息还包括纹理图像的保护带宽度字段和深度图像的保护带宽度字段,所述纹理图像的保护带宽度字段用于存储所述第i个轨道内纹理图像进行编码时采用的保护带宽度,所述深度图像的保护带宽度字段用于存储所述第i个轨道内深度图像进行编码时采用的保护带宽度。
  6. 如权利要求2所述的方法,所述第i个轨道对应的视角信息还包括第二相机标识字段和相机属性信息字段;
    所述第二相机标识字段用于存储第j个相机的标识信息;
    所述相机属性信息字段用于存储所述第j个相机的相机属性信息,所述第j个相机的相机属性信息包括以下至少一种:所述第j个相机位置的横轴分量取值、纵轴分量取值以及竖轴分量取值,所述第j个相机焦距的横轴分量取值和纵轴分量取值,以及第j个相机采集图像的分辨率宽度与高度。
  7. 如权利要求2所述的方法,所述第i个轨道对应的视角信息还包括图像信息字段,所 述图像信息字段用于存储所述第j个相机采集图像的图像信息,所述图像信息包括以下至少一个:深度图像降采样的倍数因子、深度图像左上顶点相对于平面帧原点的偏移量或纹理图像左上顶点相对于平面帧原点的偏移量。
  8. 如权利要求2所述的方法,所述第i个轨道对应的视角信息还包括自定义相机参数字段、自定义相机参数类型字段以及自定义相机参数长度字段;所述自定义相机参数字段用于存储所述第j个相机的第f个自定义相机参数,f为大于等于0且小于h的整数,h表示所述第j个相机的自定义相机参数的数量;所述自定义相机参数类型字段用于存储所述第f个自定义相机参数所属的参数类型;所述自定义相机参数长度字段用于存储所述第f个自定义相机参数的长度。
  9. 如权利要求1~8任一项所述的方法,所述方法还包括:
    获取所述沉浸媒体对应的信令描述文件,所述信令描述文件包括所述沉浸媒体对应的自由视角相机描述子,所述自由视角相机描述子用于记录每个轨道内视频片段对应的相机属性信息,所述轨道内视频片段是由所述轨道内包括的纹理图像和深度图像组成的;
    所述自由视角相机描述子被封装于所述沉浸媒体的媒体呈现描述文件的自适应集层级中,或者所述信令描述文件被封装于所述媒体呈现描述文件的表示层级中。
  10. 如权利要求9所述的方法,所述获取所述沉浸媒体的第i个轨道对应的自由视角信息数据盒,包括:
    基于所述自由视角相机描述子中记录的每个轨道内视频片段对应的相机属性信息,从所述N个相机中选择与用户所在位置信息匹配的候选相机;
    向内容制作设备发送获取来源于所述候选相机的分块视频的第一资源请求,所述第一资源请求用于指示所述内容制作设备根据所述M个轨道中每个轨道对应的自由视角信息数据盒中的视角信息,从所述M个轨道中选择第i个轨道并返回所述第i个轨道对应的自由视角信息数据盒,所述第i个轨道中封装的纹理图像和深度图像中的至少一种来自所述候选相机;
    接收所述内容制作设备返回的所述第i个轨道对应的自由视角信息数据盒。
  11. 如权利要求9所述的方法,所述获取所述沉浸媒体的第i个轨道对应的自由视角信息数据盒,包括:
    基于所述信令描述文件和用户带宽向内容制作设备发送第二资源请求,所述第二资源请求用于指示所述内容制作设备返回所述M个轨道的M个自由视角信息数据盒,一个轨道对应一个自由视角信息数据盒;
    根据每个轨道对应的自由视角信息数据盒中的视角信息和用户所在位置信息,从所述M个自由视角信息数据盒中获取第i个轨道对应的自由视角信息数据盒。
  12. 一种媒体数据处理方法,所述方法由内容制作设备执行,所述方法包括:
    将沉浸媒体封装到M个轨道中,所述沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中封装来自至少一个相机的图像,所述M个轨道属于同一个轨道组,N和M均为大于或等于1的整数;
    根据第i个轨道中w个图像的封装过程,生成第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒包括所述第i个轨道对应的视角信息;1≤i≤M,w≥1。
  13. 如权利要求12所述的方法,所述第i个轨道对应的自由视角信息数据盒包括自由视角轨道组信息,所述自由视角轨道组信息包括第一相机标识字段和图像类型字段,所述根据第i个轨道内w个图像的封装过程生成所述第i个轨道对应的自由视角信息数据盒,包括:
    确定所述w个图像来源于k个相机,k为大于1的整数;
    将所述k个相机中的第j个相机的标识信息存储在所述第一相机标识字段中;
    根据所述第j个相机采集的图像所属图像类型确定所述图像类型字段,所述图像类型包括深度图像或纹理图像中的至少一个。
  14. 如权利要求13所述的方法,所述根据所述第j个相机采集的图像所属图像类型确定所述图像类型字段,包括:
    当所述第j个相机采集的图像所属图像类型为纹理图像,则将所述图像类型字段设置为第二数值;
    当所述第j个相机采集的图像所属图像类型为深度图像,则将所述图像类型字段设置为第三数值;
    当所述第j个相机采集的图像所属图像类型为纹理图像和深度图像,则将所述图像类型字段设置为第四数值。
  15. 如权利要求12所述的方法,所述第i个轨道对应的视角信息包括视频拼接布局指示字段,所述w个图像包括纹理图像和深度图像,所述根据第i个轨道内w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:
    如果所述第i个轨道内包含的纹理图像和深度图像是拼接编码的,则将所述视频布局指示字段设置为第一数值;
    如果所述i个轨道内包含的的纹理图像和深度图像是分别编码的,则将所述视频布局指示字段设置为第二数值。
  16. 如权利要求15所述的方法,所述第i个轨道对应的视角信息还包括相机模型字段;所述根据第i个轨道内的w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:
    当采集所述w个图像的相机所属的相机模型为第一模型时,将所述相机模型字段设置为第一数值;
    当采集所述w个图像的相机所属的相机模型为第二模型时,将所述相机模型字段设置为第二数值。
  17. 如权利要求16所述的方法,所述第i个轨道对应的视角信息还包括纹理图像的保护带宽字段和深度图像的保护带宽字段,所述根据第i个轨道内w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:
    获取所述第i个轨道内在对纹理图像进行编码时采用的纹理图像的保护带宽度,并将所述纹理图像的保户带宽度存储在所述纹理图像的保护带宽度字段;
    获取所述第i个轨道内在对深度图像进行编码时采用的深度图像的保护带宽度,并将所述深度图像的保护带宽度存储在所述深度图像的保护带宽度字段。
  18. 如权利要求17所述的方法,所述第i个轨道对应的视角信息还包括第二相机标识字 段和相机属性信息字段,所述根据第i个轨道内的w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:
    将k个相机中第j个相机的标识信息存储在所述第二相机标识字段,k表示所述w个图像来源相机的数量;
    获取所述第j个相机的相机属性信息,并将获取到的相机属性信息存储在所述相机属性信息字段;
    所述第j个相机的相机属性信息包括以下任意一种或多种:所述第j个相机位置的横轴分量取值、纵轴分量取值以及竖轴分量取值,所述第j个相机焦距的横轴分量取值和纵轴分量取值,以及所述第j个相机采集图像的分辨率宽度与高度。
  19. 如权利要求18所述的方法,所述第i个轨道对应的视角信息还包括图像信息字段,所述根据第i个轨道内的w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:
    获取所述第j个相机采集的图像的图像信息,并将获取到的图像信息存储在所述图像信息字段;其中,所述图像信息包括如下一个或多个:深度图像降采样的倍数因子、深度图像左上顶点相对于平面帧原点的偏移量、纹理图像左上顶点相对于平面帧原点的偏移量。
  20. 如权利要求19所述的方法,所述第i个轨道对应的视角信息还包括自定义相机参数字段、自定义相机参数类型字段以及自定义相机参数长度字段;所述根据第i个轨道内的w个图像的封装过程生成第i个轨道对应的自由视角信息数据盒,包括:
    获取第f个自定义相机参数,并将所述第f个自定义相机参数存储在所述自定义相机参数字段,f为大于等于0且小于h的整数,h表示所述第j个相机的自定义相机参数的数量;
    确定所述第f个自定义相机参数所属的参数类型以及所述第f个自定义相机参数的长度,将所述第f个自定义相机参数所属的参数类型存储在所述自定义相机参数类型字段,以及将所述第f个自定义相机参数的长度存储在所述自定义相机参数长度字段。
  21. 如权利要求12~20任一项所述的方法,所述方法还包括:
    生成所述沉浸媒体对应的信令描述文件,所述信令描述文件包括所述沉浸媒体对应的自由视角相机描述子,所述自由视角相机描述子用于记录每个轨道内视频片段对应的相机属性信息,任一轨道内视频片段是由所述任一轨道内封装的图像组成的;
    所述自由视角相机描述子被封装于所述沉浸媒体的媒体呈现描述文件的自适应集层级中,或者所述信令描述文件被封装于所述媒体呈现描述文件的表示层级中。
  22. 如权利要求21所述的方法,所述方法还包括:
    将所述信令描述文件发送至内容消费设备,以指示所述内容消费设备基于所述自由视角相机描述子中记录的每个轨道内视频片段对应的相机属性信息,从所述多个相机中选择与用户所在位置匹配的候选相机,以及发送获取来源于所述候选相机的分块视频的第一资源请求;
    响应于所述第一资源请求,根据所述M个轨道中每个轨道对应的自由视角信息数据盒中的视角信息从所述M个轨道中选择第i个轨道并将所述i个轨道对应的自由视角信息数据盒发送至所述内容消费设备。
  23. 如权利要求21所述的方法,所述方法还包括:
    将所述信令描述文件发送至内容消费设备,以指示所述内容消费设备根据所述信令描述文件和用户带宽发送第二资源请求;
    响应于所述第二资源请求,将所述M个轨道对应的M个自由视角信息数据盒发送至所述内容消费设备,以指示所述内容消费设备根据每个轨道对应的自由视角信息数据盒中的视角信息和用户所在位置信息,从所述M个自由视角信息数据盒中获取第i个轨道对应的自由视角信息数据盒。
  24. 一种沉浸媒体的数据处理装置,所述沉浸媒体是由处于不同视角的N个相机拍摄的图像组成,所述沉浸媒体被封装到M个轨道,一个轨道中封装来自至少一个相机的图像,N和M均为大于1的整数;所述装置包括:
    获取单元,用于获取所述沉浸媒体的第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒包括所述第i个轨道对应的视角信息,i为大于或等于1,且小于或等于M的整数;
    处理单元,用于根据所述自由视角信息数据盒中的视角信息对所述第i个轨道内封装的图像进行解码。
  25. 一种沉浸媒体的数据处理装置,包括:
    封装单元,用于将所述沉浸媒体封装到M个轨道中,所述沉浸媒体是由处于不同视角的N个相机拍摄的图像组成的,一个轨道中封装来自至少一个相机的图像,所述M个轨道属于同一个轨道组,N和M均为大于或等于1的整数;
    生成单元,用于根据第i个轨道中w个图像的封装过程,生成第i个轨道对应的自由视角信息数据盒,所述自由视角信息数据盒包括所述第i个轨道对应的视角信息;1≤i≤M,w≥1。
  26. 一种计算机设备,包括:
    处理器,适于实现一条或多条计算机程序;以及
    计算机存储介质,所述计算机存储介质存储有一条或多条计算机程序,所述一条或计算机程序程序适于由处理器加载并执行如权利要求1~23任一项所述的方法。
  27. 一种计算机存储介质,所述计算机存储介质中存储有第一计算机程序程序和第二计算机程序,所述第一计算机程序被处理器执行时,用于执行如权利要求1-11任一项所述的沉浸媒体的数据处理方法;所述第二计算机程序被处理器执行时,用于执行如权利要求12-23任一项所述的沉浸媒体的数据处理方法。
  28. 一种包括指令的计算机程序产品,当其在计算机上运行时,使得所述计算机执行权利要求1~23任一项所述的方法。
PCT/CN2022/080257 2021-06-11 2022-03-11 沉浸媒体的数据处理方法、装置、相关设备及存储介质 Ceased WO2022257518A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US17/993,852 US12395615B2 (en) 2021-06-11 2022-11-23 Data processing method and apparatus for immersive media, related device, and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202110659190.8A CN115474034B (zh) 2021-06-11 2021-06-11 沉浸媒体的数据处理方法、装置、相关设备及存储介质
CN202110659190.8 2021-06-11

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US17/993,852 Continuation US12395615B2 (en) 2021-06-11 2022-11-23 Data processing method and apparatus for immersive media, related device, and storage medium

Publications (1)

Publication Number Publication Date
WO2022257518A1 true WO2022257518A1 (zh) 2022-12-15

Family

ID=84363490

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2022/080257 Ceased WO2022257518A1 (zh) 2021-06-11 2022-03-11 沉浸媒体的数据处理方法、装置、相关设备及存储介质

Country Status (4)

Country Link
US (1) US12395615B2 (zh)
CN (1) CN115474034B (zh)
TW (1) TWI796989B (zh)
WO (1) WO2022257518A1 (zh)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN117278820A (zh) * 2023-09-18 2023-12-22 腾讯科技(深圳)有限公司 视频生成方法、装置、设备及存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20100128112A1 (en) * 2008-11-26 2010-05-27 Samsung Electronics Co., Ltd Immersive display system for interacting with three-dimensional content
WO2019066191A1 (ko) * 2017-09-28 2019-04-04 엘지전자 주식회사 스티칭 및 리프로젝션 관련 메타데이터를 이용한 6dof 비디오를 송수신하는 방법 및 그 장치
CN111264058A (zh) * 2017-09-15 2020-06-09 交互数字Vc控股公司 用于对三自由度和体积兼容视频流进行编码和解码的方法、设备
CN112492289A (zh) * 2020-06-23 2021-03-12 中兴通讯股份有限公司 沉浸媒体数据的处理方法及装置、存储介质和电子装置
CN112804256A (zh) * 2021-02-09 2021-05-14 腾讯科技(深圳)有限公司 多媒体文件中轨道数据的处理方法、装置、介质及设备

Family Cites Families (32)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR100446635B1 (ko) * 2001-11-27 2004-09-04 삼성전자주식회사 깊이 이미지 기반 3차원 객체 표현 장치 및 방법
US9742974B2 (en) * 2013-08-10 2017-08-22 Hai Yu Local positioning and motion estimation based camera viewing system and methods
KR20160112898A (ko) * 2015-03-20 2016-09-28 한국과학기술원 증강현실 기반 동적 서비스 제공 방법 및 장치
WO2017024990A1 (en) 2015-08-07 2017-02-16 Mediatek Inc. Method and apparatus of bitstream random access and synchronization for multimedia applications
US10217189B2 (en) * 2015-09-16 2019-02-26 Google Llc General spherical capture methods
US10820025B2 (en) 2016-02-03 2020-10-27 Mediatek, Inc. Method and system of push-template and URL list for dash on full-duplex protocols
US10404411B2 (en) 2016-02-19 2019-09-03 Mediatek Inc. Method and system of adaptive application layer FEC for MPEG media transport
GB2553315A (en) * 2016-09-01 2018-03-07 Nokia Technologies Oy Determining inter-view prediction areas
CN109716759B (zh) 2016-09-02 2021-10-01 联发科技股份有限公司 提升质量递送及合成处理
US20180075576A1 (en) 2016-09-09 2018-03-15 Mediatek Inc. Packing projected omnidirectional videos
US10623635B2 (en) 2016-09-23 2020-04-14 Mediatek Inc. System and method for specifying, signaling and using coding-independent code points in processing media contents from multiple media sources
US11197040B2 (en) 2016-10-17 2021-12-07 Mediatek Inc. Deriving and signaling a region or viewport in streaming media
US10742999B2 (en) 2017-01-06 2020-08-11 Mediatek Inc. Methods and apparatus for signaling viewports and regions of interest
US10805620B2 (en) 2017-01-11 2020-10-13 Mediatek Inc. Method and apparatus for deriving composite tracks
US11139000B2 (en) 2017-03-07 2021-10-05 Mediatek Inc. Method and apparatus for signaling spatial region information
US10542297B2 (en) 2017-03-07 2020-01-21 Mediatek Inc. Methods and apparatus for signaling asset change information for media content
US10778993B2 (en) * 2017-06-23 2020-09-15 Mediatek Inc. Methods and apparatus for deriving composite tracks with track grouping
US10565616B2 (en) 2017-07-13 2020-02-18 Misapplied Sciences, Inc. Multi-view advertising system and method
EP3777137B1 (en) * 2018-04-06 2024-10-09 Nokia Technologies Oy Method and apparatus for signaling of viewing extents and viewing space for omnidirectional content
US11146802B2 (en) * 2018-04-12 2021-10-12 Mediatek Singapore Pte. Ltd. Methods and apparatus for providing two-dimensional spatial relationships
WO2020013484A1 (ko) * 2018-07-11 2020-01-16 엘지전자 주식회사 360 비디오 시스템에서 오버레이 처리 방법 및 그 장치
JP7271099B2 (ja) * 2018-07-19 2023-05-11 キヤノン株式会社 ファイルの生成装置およびファイルに基づく映像の生成装置
US11178373B2 (en) * 2018-07-31 2021-11-16 Intel Corporation Adaptive resolution of point cloud and viewpoint prediction for video streaming in computing environments
WO2020091404A1 (ko) * 2018-10-30 2020-05-07 엘지전자 주식회사 비디오 송신 방법, 비디오 전송 장치, 비디오 수신 방법 및 비디오 수신 장치
WO2020189895A1 (ko) 2019-03-21 2020-09-24 엘지전자 주식회사 포인트 클라우드 데이터 송신 장치, 포인트 클라우드 데이터 송신 방법, 포인트 클라우드 데이터 수신 장치 및 포인트 클라우드 데이터 수신 방법
US11308925B2 (en) * 2019-05-13 2022-04-19 Paul Senn System and method for creating a sensory experience by merging biometric data with user-provided content
US11276142B2 (en) * 2019-06-27 2022-03-15 Electronics And Telecommunications Research Institute Apparatus and method for synthesizing virtual viewpoint images
US11831861B2 (en) 2019-08-12 2023-11-28 Intel Corporation Methods for viewport-dependent adaptive streaming of point cloud content
US11729243B2 (en) * 2019-09-20 2023-08-15 Intel Corporation Dash-based streaming of point cloud content based on recommended viewports
US11315289B2 (en) * 2019-09-30 2022-04-26 Nokia Technologies Oy Adaptive depth guard band
US11804042B1 (en) * 2020-09-04 2023-10-31 Scale AI, Inc. Prelabeling of bounding boxes in video frames
US11489899B1 (en) * 2021-04-12 2022-11-01 Comcast Cable Communications, Llc Segment ladder transitioning in adaptive streaming

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20100128112A1 (en) * 2008-11-26 2010-05-27 Samsung Electronics Co., Ltd Immersive display system for interacting with three-dimensional content
CN111264058A (zh) * 2017-09-15 2020-06-09 交互数字Vc控股公司 用于对三自由度和体积兼容视频流进行编码和解码的方法、设备
WO2019066191A1 (ko) * 2017-09-28 2019-04-04 엘지전자 주식회사 스티칭 및 리프로젝션 관련 메타데이터를 이용한 6dof 비디오를 송수신하는 방법 및 그 장치
CN112492289A (zh) * 2020-06-23 2021-03-12 中兴通讯股份有限公司 沉浸媒体数据的处理方法及装置、存储介质和电子装置
CN112804256A (zh) * 2021-02-09 2021-05-14 腾讯科技(深圳)有限公司 多媒体文件中轨道数据的处理方法、装置、介质及设备

Also Published As

Publication number Publication date
TWI796989B (zh) 2023-03-21
TW202249496A (zh) 2022-12-16
CN115474034B (zh) 2024-04-26
US20230088144A1 (en) 2023-03-23
US12395615B2 (en) 2025-08-19
CN115474034A (zh) 2022-12-13

Similar Documents

Publication Publication Date Title
TW201840178A (zh) 適應性擾動立方體之地圖投影
CN113852829A (zh) 点云媒体文件的封装与解封装方法、装置及存储介质
US12010402B2 (en) Data processing for immersive media
CN115396647A (zh) 一种沉浸媒体的数据处理方法、装置、设备及存储介质
CN116781676A (zh) 一种点云媒体的数据处理方法、装置、设备及介质
CN113949829B (zh) 媒体文件封装及解封装方法、装置、设备及存储介质
TWI796989B (zh) 沉浸媒體的數據處理方法、裝置、相關設備及儲存媒介
CN116456166A (zh) 媒体数据的数据处理方法及相关设备
WO2022116822A1 (zh) 沉浸式媒体的数据处理方法、装置和计算机可读存储介质
CN113766272B (zh) 一种沉浸媒体的数据处理方法
US12489927B2 (en) File decapsulation method and apparatus for free viewpoint video, device, and storage medium
CN113497928B (zh) 一种沉浸媒体的数据处理方法及相关设备
HK40074454B (zh) 一种沉浸媒体的数据处理方法及设备
HK40089839A (zh) 媒体数据的数据处理方法及相关设备
HK40084577B (zh) 自由视角视频的文件封装方法、装置、设备及存储介质
HK40074454A (zh) 一种沉浸媒体的数据处理方法及设备
HK40087291A (zh) 一种沉浸媒体的数据处理方法及相关装置
HK40072001A (zh) 沉浸式媒体的数据处理方法、装置和计算机可读存储介质
HK40091103A (zh) 一种沉浸媒体的数据处理方法、装置、设备及存储介质
CN115102932A (zh) 点云媒体的数据处理方法、装置、设备、存储介质及产品
HK40086079B (zh) 点云媒体文件封装方法、装置、设备及存储介质
HK40086079A (zh) 点云媒体文件封装方法、装置、设备及存储介质
CN116137664A (zh) 点云媒体文件封装方法、装置、设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22819123

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 11202305312V

Country of ref document: SG

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 08.05.2024)

122 Ep: pct application non-entry in european phase

Ref document number: 22819123

Country of ref document: EP

Kind code of ref document: A1